Most connected TV (CTV) incrementality tests fail before a single ad runs. They don’t fail because the media underperformed, but because the test produced a number that looks convincing without proving anything.
This failure matters more every quarter because the people who control the budget increasingly want more proof that CTV works. Coming into this year, cross-platform measurement ranked as a top priority for 72% of advertisers, up from 64% in 2025, according to the IAB’s 2026 Outlook Study. Advertisers need to prove that investment in streaming delivers results that search and social can’t achieve.
The most honest way to address this challenge is through an incrementality test. It involves holding back a control group, advertising to the remaining audience, and measuring the difference in outcomes. While the concept is straightforward, execution can be challenging. Even the largest, best-resourced advertising teams have struggled to answer the incrementality question. If the biggest spenders find it difficult, the rest of us need a reliable method.
Here’s how to design a CTV incrementality test for B2B that holds up, avoid the mistakes that quietly invalidate most incrementality tests, and turn the result into a number your CFO will accept.
10X your SEO with Semrush for Enterprise.
The world’s most powerful SEO platform, purpose-built for Enterprise.
Request demo
Incrementality isn’t attribution
Attribution focuses on identifying which advertising exposure deserves credit for a conversion. In contrast, incrementality tackles a more complex question: Would that conversion have happened without the ad? The distinction between these two lines of inquiry highlights the shortcomings of CTV measurement.
The most common incrementality testing method compares people who saw the ad with those who didn’t, then reports the difference as a measure of uplift. While this may seem rigorous, it’s a flawed approach. People who viewed your CTV ad are inherently different from those who didn’t. They tend to stream more, may already be interested in your product, and are deliberately targeted. This measurement captures the effects of targeting rather than the impact of the ad itself. It’s like attributing someone’s fitness to their gym attendance simply because they chose to go.
Real incrementality requires deciding, before a single impression runs, which accounts you deliberately won’t advertise to. That withheld group is the control. What happens to them is what would have happened without exposure to your advertising. Everything the exposed group does above that baseline is incremental. If you skip this step, nothing you do later in the analysis can make up for it.
Pick the right unit to randomize
Unlike many other forms of digital marketing, CTV has no persistent cookie. Ads are served to households and IP addresses, and several people can watch the same screen at the same time, so you can’t cleanly hold an individual-level split. Go up a level instead.
B2B still targets individuals across channels (e.g., the named members of the buying committee, on LinkedIn, in email, and in your ABM program), but that committee makes the purchase decision as a group, and ABM already reports at the account level. A complex B2B deal involves a group of decision-makers, not just one. Randomize where the decision lands. The account is the correct unit of randomization.
For account-based programs, divide your target account list into two groups. Half the accounts are eligible for CTV advertising, while the other half are excluded from all advertising, including streaming. This approach aligns with how B2B reporting is typically structured, focusing on accounts and their progression through the sales pipeline.
Choose the testing method and understand what it proves
Not all testing methods are equal. From the strongest proof to the weakest:
- An account-matched holdout with hard suppression is the strongest test most B2B teams can run. Exclude the control accounts’ IP addresses and device IDs from every related line item, not just the test campaign.
- A geo-matched market test uses a difference-in-differences design, highlighting the change in the test markets minus the change in the control markets over the same period.
- Media mix modeling is a fallback, not a test. It’s a correlation between spend and outcomes. It’s useful for planning, but don’t dress it up as causal proof of a single channel.
Select the metric before you begin testing, and avoid focusing on early-stage metrics like impressions because they indicate delivery rather than impact.
For B2B, the primary metric should be qualified pipeline or opportunities generated from target accounts. You can also use faster signals, such as target account site visits, increases in branded search, improved paid social metrics, and demo requests, to gauge whether the test is working before the pipeline has time to mature.
Ensure your test size is statistically significant
The primary reason most incrementality tests fail isn’t the media. B2B conversions are both rare and slow. If your baseline opportunity rate is low and you hold back only a small number of accounts, you won’t have enough conversions to detect any meaningful improvement. As a result, the test will indicate “no effect,” regardless of how well CTV performed.
Before committing, check whether the test can reach statistical significance at all. Three things decide this: how often these accounts convert today, how large a lift you’d need to call the result real, and how many accounts you can put on each side of the split.
Unfortunately, when conversions are rare and the account list is short, no sample will surface an effect, even if CTV is working. If that’s where you land, you have two choices: Make an upper-funnel metric the primary one and treat pipeline as directional, or don’t run the test yet. A test that’s too small to detect an effect is worse than no test at all because it produces a confident wrong answer.
Protect the control group
Keep two things in mind when conducting incrementality tests. First, measure both the test and control groups before the campaign starts to ensure they’re on parallel trends. If the control group already performed better, any perceived lift from the campaign may be misleading.
Second, protect against data leakage. Factors such as shared corporate IP addresses, co-viewing, and users exposed to multiple active campaigns within the same account can contaminate your control group. Enforce suppression across all line items, and monitor it closely during the test.
Also, allow enough time for the campaign to take effect. Upper-funnel signals typically take weeks to emerge, whereas the sales pipeline requires the campaign duration plus an appropriate lookback period aligned with your sales cycle. If you measure a two-week test against pipeline, it may seem like a failure simply because the pipeline hasn’t had enough time to develop.
Why disciplined teams keep their budget
A clean incrementality test only matters if the result lands with the person who controls the budget. Your CFO doesn’t think in lift percentages. They think in pipeline, cost per opportunity, and how long the spend takes to pay for itself. Report incremental pipeline, cost per incremental opportunity, and payback period. Those are the numbers you can defend in a budget meeting.
The teams that keep their CTV budget through 2026 won’t be the ones with the best-looking dashboards. They’ll be the ones who decided — before the campaign ever launched — which accounts they were willing to leave alone. None of that is glamorous. It’s the difference between knowing CTV worked and hoping it did. Build the control group first. The rest follows.