Building a Waterfall Enrichment Stack in Clay for Maximum Match Rates

Sequencing cheap providers first cuts costs while boosting match rates.

Senior Staff Writer · · 10 min read
Cover illustration for “Building a Waterfall Enrichment Stack in Clay for Maximum Match Rates”
B2B Data and Enrichment · October 6, 2026 · 10 min read · 2,252 words

Building a waterfall enrichment stack in Clay is a specific, learnable construction process, not a vague best practice. Provider selection, sequencing logic, and field-level configuration compound match rates far beyond what any single data source can deliver, and this piece walks through how each piece is built, in order, starting with the structural reason the whole approach exists.

Why Single Providers Fail at Quality and Coverage

No enrichment provider wins on both quality and coverage at once, because the two work against each other by design. A vendor that checks each record against a primary source returns data worth trusting, but only for the records it has actually verified, so its coverage shrinks. A vendor that pulls from many databases at once reaches nearly any account on the list, but ships records that are stale or simply wrong. Picking only one provider decides the failure mode in advance: a clean list that is half empty, or a full list nobody should rely on.

This tradeoff is baked into how verification and aggregation work as methods, so chaining providers together solves the problem that switching providers never will. Email and phone fields show this most clearly, since no single provider holds strong global coverage and match rates swing hard depending on region and seniority. A provider that performs well in one region or for one seniority band tends to cover exactly where another provider falls short, so running them in sequence lets coverage build on itself across the chain. That compounding effect, not any single vendor's strength, is the reason a waterfall is the correct architecture for this problem.

How a waterfall works mechanically before any configuration decisions are made

A waterfall is a fallback chain. Records move through a fixed order of providers and exit the chain the moment one of them returns a confident result. A record never reaches the second or third provider unless the one above it came back empty, and that single rule shapes the entire cost structure.

Because the chain stops at the first hit, credits get spent only on the lookup that actually finds something. Clay's own waterfall guide shows this in practice: an Infer Email step runs first, for free, building a likely address out of a name and a domain before any paid provider gets touched. Credits only get spent once a paid provider lands a verified match. A contact found at the first paid step never reaches the second, so the second provider's credit never gets charged for it. Only the genuinely hard records, the ones the cheap broad sources could not resolve, ever make it down to the expensive specialist waiting at the bottom of the chain.

This one property, sequential movement with an early stop, is what makes the cost model and the coverage math work in the same direction at once, so running it right makes the stack cheaper and more complete at the same time. The logic holds regardless of what field is being filled in. It applies to work emails, mobile numbers, firmographic details, and profile URLs alike. The data type changes from one waterfall to the next, but the chain logic stays the same.

The pre-enrichment gate: filtering your list before the waterfall runs

The single biggest way practitioners burn through credits is running enrichment against an unfiltered list. Enriching rows that would never pass an ICP check wastes a large share of a credit budget before the waterfall has done anything useful at all, since every one of those lookups still costs money whether the resulting contact is usable or not.

The correct order runs the other direction from how most people set it up: qualify the account first, find the contacts at that account second, and only then run contact enrichment. A useful way to build this in practice is to import companies first, then use a "find people" step to expand each already-qualified account into its decision-makers, so that contact enrichment only fires on people who sit inside accounts that have already cleared the bar.

Two more checks protect credits once the waterfall itself is running. A conditional check that skips enrichment when a verified email already sits in the table prevents Clay from re-running a lookup it does not need to, which can save a meaningful share of credits on any table that already has partial data filled in. A max-credit cap set on a per-row basis stops a single hard-to-find record from working its way through the entire chain and eating far more credit than that record is worth. A contact that expensive to track down is probably outside the ICP to begin with.

None of this works, though, if the underlying company data is messy. Name and URL normalization has to happen before any provider query runs, because enrichment APIs match literally: "Vanderbuild Strategy Group, LLC" and "Vanderbuild" can fail to match on the very same record if the strings are not cleaned first. Legal suffixes like LLC, Ltd, and GmbH need to be stripped and root domains cleaned before the first enrichment step ever fires. None of this is optional hygiene tacked onto the real work. Skipping it makes the waterfall fight against itself, burning credits on mismatches that better formatting would have avoided. A properly filtered, properly normalized table is the precondition that lets the sequencing logic in the next section actually function the way it is designed to.

Sequencing logic: ordering providers by cost-then-coverage, not by brand reputation

Diagram: Why Provider Order Changes Everything: The Cost Math. Visualizes: Visualize the concrete cost comparison between a correctly sequenced waterfall and one that runs the expensive provider first.

Order is the only real decision inside a waterfall. The cheapest provider that returns trustworthy data for the ICP in question goes first, and every fallback after it exists only to catch what the level above it missed. Sequencing by brand name or by a provider's overall reputation for quality produces a worse result than sequencing by cost-per-confident-hit, because reputation says nothing about where a given provider actually performs well against a given list.

Clay's enrichment guide works through an example that makes the math concrete. With Provider A charging 1 credit per hit, Provider B charging 2, and Provider C charging 4, a correctly ordered waterfall reaches an 82% match rate at roughly 1.77 credits per valid email. Running the expensive provider first instead means that single step alone can cost more than the entire correctly sequenced waterfall, while still landing a lower match rate overall. The order itself is the entire difference between an efficient stack and a wasteful one.

Provider placement should also bend to ICP geography and persona, not price alone. A provider that handles European contacts well might sit lower in the chain for a list of U.S. accounts, and the right position for any given provider depends on where it actually performs, not on its general reputation in the market. A free inference step runs first at zero cost, building a likely address from name and domain. Cheap, decent-coverage providers then clear the bulk of the list at the lowest price per record. Expensive specialists run last, touching only the small number of records nobody else could resolve. Clay's waterfall itself cascades across more than 100 email providers in this kind of sequence, charging credits only for the one that actually returns a match, with the free Infer Email step running before any paid provider gets touched. Running the premium provider first means the stack pays full price to enrich records a free inference step could have handled on its own. That wasted cost compounds across every record the expensive provider touches before the cheaper fallbacks ever get a chance to run.

Building the email waterfall: the stack that carries most outbound programs

Diagram: The Email Waterfall Stack: Free to Fallback. Visualizes: Show the four-stage email enrichment waterfall in sequence, as a vertical stepped flow.

Email carries most outbound programs, and the email waterfall is the stack most practitioners build first, since no single provider covers both the volume and the accuracy a working program actually needs. A common sequencing pattern used in production starts with the free inference step, building a likely address from name and domain before any paid provider is touched. From there, a high-accuracy primary source, something like LeadMagic or Findymail, handles the bulk of the list at the lowest paid cost per hit. A second provider with a different underlying database, Findymail again in some configurations, catches the next layer of misses that the first paid step could not resolve. A final fallback, such as Datagma or Hunter, picks up whatever remains, and because this step touches the fewest records, it can justify a noticeably higher per-credit cost without blowing up the budget for the stack as a whole.

Clay's 2025 work email benchmark found that Hunter returns the most trustworthy emails of the providers tested, but reaches barely half of contacts on its own. That exact combination, high trust and low coverage, is why Hunter belongs at the bottom of the chain as a high-accuracy fallback. Once the waterfall has actually found an address, an email validation step, using a tool like ZeroBounce or NeverBounce, or Clay's own built-in validation, should run before the record leaves the table. Deliverability protection belongs inside the waterfall itself, built into the workflow, rather than as a separate check bolted on right before a campaign goes out.

Pedro Azevedo's work at Coverflex shows what this design is actually for. "We built an automated flow that identifies signals, enriches data, and pushes leads to our sales team only when it's most relevant. Instead of flooding HubSpot with data, we only push leads once they've shown strong signals. This helps our sales team stay focused and reduces noise." That is the entire point of sequencing a waterfall carefully: not raw volume into a CRM, but a filtered stream of contacts that a sales team can actually act on.

Building the phone waterfall: a distinct and more expensive enrichment problem

Phone enrichment is not a copy of the email stack with different providers swapped in. It carries its own economics, its own provider choices, and its own conditions for when it makes sense to build. A phone waterfall fits cold calling teams working high-value accounts where a direct dial actually changes the outcome of the outreach. It is not something to bolt onto every outbound table as a default.

For C-suite targets specifically, the provider landscape splits along fairly distinct lines. One provider shows strong hit rates on personal emails sourced from a professional social network. Lusha covers US-based direct dials well, though fill rates drop for C-level targets specifically. Datagma finds mobile numbers through social profiles, and Fullenrich queries across multiple databases at once. The sequencing logic matches the email waterfall exactly: the cheapest confident hit goes first, and the expensive specialist runs last. What changes is the tolerance for error. Phone lookups run roughly eight times the cost of email lookups, so the credit math demands much tighter ICP filtering before the phone waterfall runs at all, since a failed lookup at phone rates burns through budget far faster than a failed lookup on email ever could. A phone waterfall built without tight filtering upstream is one of the more reliable ways to exhaust a credit budget on records that were never going to convert.

Building the company enrichment waterfall: the gate that runs before contact work

Company enrichment is the gate that decides which accounts are worth spending any contact-enrichment credit on in the first place, and it has to run before either the email or phone waterfall touches a single record.

The chaining principle at work here carries through the entire stack. A provider like Clearbit returns a clean company domain, that domain feeds directly into the email-finder waterfall, and the verified email that comes out the other end feeds into a deliverability check before anything gets pushed downstream. Each provider in the chain hands the next one a better input than the raw list started with, which is the same compounding logic described earlier in the piece, just applied one level up, at the company rather than the contact. A clean first build, for practitioners setting this up for the first time, comes down to five pieces: one source list, a company enrichment gate, a contact waterfall, an ICP filter, and one destination. Everything beyond that basic loop is an upgrade worth adding only once this core sequence is already producing clean, reliable data on its own.

Credit management under Clay's current pricing structure

Clay's pricing overhaul, which took effect on March 11, 2026, changed the cost model in ways that directly affect how a waterfall should be sized and watched day to day. A 30% surcharge now applies to mid-month credit top-ups above the plan rate, down from a higher surcharge rate that applied before the overhaul, which still makes running over budget meaningfully more expensive than staying within it. Understanding the new structure matters for anyone building or repricing a stack under the current system.

Credits themselves are now split into two categories: Data Credits, which pay for the enrichment data a waterfall actually pulls in, and Actions, which pay for using the platform itself. Data costs dropped by a substantial margin across most enrichment providers under the new structure. The old plan tiers, Starter, Explorer, and Pro, have been replaced by Free, Launch, and Growth plans, alongside a custom Enterprise tier for larger teams. For anyone sizing a waterfall against a real budget, that split between Data Credits and Actions is the number to build around: it determines how far a given plan actually stretches once the filtering, sequencing, and field-level waterfalls described above are all running at once.

More in B2B Data and Enrichment