The 5 Pillars of Incrementality Testing With AdBeacon: A Framework for Proving True Ad Profit
Incrementality testing has crossed from niche practice to mainstream expectation faster than most measurement trends do. Just two years ago it was a technique reserved for data science teams at the biggest brands.
Now, 52 percent of US brand and agency marketers run incrementality tests, and in retail media specifically, 71 percent of advertisers rank incrementality as their single most important KPI.
The reason isn’t hard to find once you look at the gap it keeps uncovering: platform-reported ROAS is running 20 to 60 percent above measured incremental lift across the accounts being tested.
That gap is exactly what incrementality testing pillars need to be built to close, not paper over. This framework lays out the five pillars AdBeacon treats as non-negotiable for a test that actually holds up once you act on it.
Incrementality vs Attribution: Why the Distinction Matters
Attribution assigns credit for a conversion to whichever touchpoint a model decides deserves it, whether that’s the last click, the first click, or something in between.
Incrementality asks a different, causal question: would this sale have happened without the ad at all? Attribution tells you who to thank. Incrementality tells you who to keep paying. Both matter, but only one of them can tell you whether a channel getting credit for a sale actually earned it.
The 5 Pillars of AdBeacon’s Incrementality Testing Framework
Pillar 1: First-Party, Click-Only Signal as the Foundation
- An incrementality test is only as trustworthy as the conversion data feeding it. Running a holdout test on top of platform-reported, view-through-inflated numbers just produces a more elaborate version of the same inflated answer.
- AdBeacon’s incrementality testing framework starts from verified, first-party, click-only conversion data, the same ground-truth signal the rest of AdBeacon’s reporting is built on, so the test and the baseline you’re comparing it to are speaking the same language.
Pillar 2: Holdout or Geo Testing to Establish Causation
This is the actual mechanism.
- You withhold a channel or campaign from a matched audience or region, a market, an audience segment, a percentage of eligible users, and compare what happens against a group that still sees the ads.
- The delta between the two is your incremental lift. One recent case example: a brand running a geo holdout on Meta saw an 18 percent sales lift in exposed markets against a matched control, implying an incremental ROAS of 1.9, well below the 4.2 the platform had reported. That’s the kind of gap this pillar exists to catch.
Pillar 3: A Clean Pre-Test Baseline
- You can’t measure lift against a number you never established. Before a holdout starts, you need a clear read on what performance looked like without the test running, ideally over a period that doesn’t overlap with a major sale, a seasonal spike, or a pricing change.
- Skip this step and a real lift can look like noise, or noise can look like a real lift, because there’s nothing stable to compare either one against.
Pillar 4: Statistical Rigor, Sample Size and Test Duration
A test that ends too early or runs on too small a group produces a number that feels like an answer but isn’t one.
- Holdout size and test duration trade off against each other directly: a smaller holdout collects less signal per day, so it needs to run longer to reach statistical significance, while a larger holdout reaches significance faster but costs more in suppressed conversions during the test.
- Most digital holdouts need somewhere between two and six weeks to reach a reliable read, longer for upper-funnel channels where conversions are sparser. Pre-registering the test, meaning fixing the start date, end date, and primary metric before launch, matters too.
- Checking results early and stopping the moment they look favorable inflates false positives and makes the final number meaningless.
Pillar 5: Ongoing Re-Testing, Not a One-Time Exercise
Incrementality decays.
- A channel that showed strong lift last quarter can show something very different once creative fatigues, a competitor changes their bidding, or the season shifts consumer behavior entirely.
- Teams that treat a single test as a permanent verdict tend to keep spending on the strength of a result that quietly stopped being true months earlier.
- The teams that retest on a defined cadence, commonly quarterly for the channels carrying the most budget, catch that drift before it compounds across a full planning cycle.
- This is also where a result stops being a one-off finding and starts feeding back into a broader measurement system, calibrating an MMM like Meridian rather than sitting in a slide deck until someone remembers to look at it again.
How to Run Your First Test Using This Framework
Start with the channel or tactic you trust least, usually branded search or retargeting, the two categories attribution over-credits most consistently because the customer was often going to convert anyway.
Size your holdout at roughly 15 to 20 percent of the eligible audience, run it for a minimum of four weeks, and hold everything else about the campaign steady for the duration.
When it’s done, feed the result back into your blended reporting rather than filing it away, and put the same channel back on the calendar for a retest next quarter.
The value of a framework like this isn’t the test itself. It’s what a brand does with a number it can actually trust. If you want help designing your first holdout test against verified, first-party data, book a live AdBeacon demo and we’ll walk through what it looks like on your own account.
—-
FAQ
What’s the difference between incrementality testing and attribution?
Attribution assigns credit for a conversion to a touchpoint using a model, first click, last click, or a blend. Incrementality testing uses a controlled experiment, typically a holdout, to answer a causal question: would that sale have happened without the ad. Attribution can overstate a channel’s value; incrementality is what confirms or corrects it.
How long should an incrementality test run?
Most digital holdout tests need two to six weeks to reach statistical significance, depending on holdout size and conversion volume. Upper-funnel channels with sparser direct conversions, like CTV or podcast advertising, often need four to six weeks or longer.
How big should my holdout group be?
Fifteen to twenty percent of the eligible audience is a common starting point. A larger holdout, such as a 50/50 split, reaches statistical significance faster but suppresses more potential conversions during the test window, so the right size depends on how quickly you need an answer versus how much opportunity cost you can absorb.
How often should I retest incrementality?
Quarterly is the standard cadence for the channels carrying the most budget. Incrementality isn’t fixed. Creative fatigue, competitor behavior, and seasonality all shift what a channel is actually driving, so a result from two quarters ago may no longer reflect what’s true today.
Do I need incrementality testing if I already use multi-touch attribution?
Yes. Multi-touch attribution improves on last-click by crediting more of the customer journey, but it’s still a model, not a causal test. Incrementality testing is what validates whether the credit a model assigns actually reflects a sale the ad caused, and it’s the layer that catches channels an attribution model consistently over-credits.
Sources
- Eightx: What Is Incrementality Testing, A 2026 Guide for DTC Operators
- AdSights: Incrementality Testing, Definition and Examples
- Haus: How Long Should You Run an Incrementality Test For
- Haus: 3 Reasons to Retest Your Incrementality Experiments
- Measured: Incrementality Measurement Is a Program, Not a Project
- fusepoint: Incrementality Testing, How to Measure Incremental Lift and Optimize Spend
- Coupler.io: Ecommerce Analytics Trends in 2026