Key Takeaways and Introduction
- TV and CTV break click-based attribution because most conversions happen on a different device than the ad exposure, which forces platforms to rely on inflated view-through and device-graph matching.
- There are four methodology families for measuring TV and CTV incrementality: geo holdout tests, platform lift studies, panel and ACR-based attribution, and media mix modeling (MMM). Only geo experiments produce ground-truth incrementality without relying on user-level tracking.
- Geo holdout testing measures TV and CTV impact from transaction data alone, which makes it immune to attribution windows, cookie deprecation, and platform privacy changes.
- Measurement is where the money is: Measured’s Q1 2026 analysis of 148 video advertisers found CTV delivers the highest median incremental ROAS among video channels ($1.38), but the middle 50% of advertisers range from $0.50 to $3.07. Brands that measure incrementality and allocate accordingly capture the top of that range.
- In one Measured geo holdout on streaming TV, roughly half of the conversions the ad platform claimed were actually incremental, and the channel delivered a $2.46 incremental ROAS at 90% statistical significance.
- The strongest measurement programs triangulate: geo experiments produce causal ground truth, and those results calibrate an always-on MMM so every week of TV and CTV spend gets an incrementality-adjusted read.
US connected TV ad spend is projected to reach roughly $38 billion in 2026, growing about 14% year over year, and for the first time CTV upfront commitments are expected to exceed primetime linear TV upfronts (eMarketer, 2026). The money has moved. The measurement, for most brands, has not.
Measured’s own Q1 2026 analysis of 148 brands investing in video makes the stakes concrete: CTV delivers the highest median incremental ROAS of any major video channel at $1.38, ahead of YouTube at $1.26 and Linear TV at $0.75. But the middle 50% of advertisers range all the way from $0.50 to $3.07. A spread that wide is not the channel behaving inconsistently. It is the difference between brands that measure and structure CTV deliberately and brands that take platform reporting at face value. (The full analysis, including how enterprise brands more than double the median return, is in the report CTV is Mainstream. Optimization Just Hasn’t Caught Up.)
That spread is what this guide is about. Landing in the top half of it starts with knowing what your TV and CTV spend actually contributes, so this article compares the four methodologies for measuring TV and CTV incrementality, explains how each one works, and shows what a real geo holdout test on streaming TV looks like in practice.
If you are new to incrementality as a concept, start with What is Incrementality in Marketing and come back. This article assumes you know why platform-reported numbers overstate impact and focuses on how to get a true read specifically for television.
Why TV and CTV Break Click-Based Measurement
Digital attribution was built on a simple chain: ad impression, click, conversion, all on the same device, all trackable. TV and CTV break every link in that chain.
There is no click. A viewer sees a streaming ad on the living room TV and later converts on a phone or laptop. The platform never observes a click path, so it falls back on view-through attribution: if a device in the household saw the ad and any device in the household converted within the attribution window, the platform takes credit. View-through windows of 7 to 30 days mean CTV platforms routinely claim conversions that were going to happen anyway.
Device graphs are probabilistic. Connecting the TV that showed the ad to the phone that converted requires identity resolution across devices, usually through IP matching. IP-based matching degrades as privacy features proliferate, and it was never deterministic to begin with.
Walled gardens grade their own homework. YouTube, Amazon, Roku, Hulu, and Netflix each measure their own performance with their own attribution logic, their own windows, and no common standard. Summing their reported conversions routinely produces a number larger than your actual business.
Linear TV has no user-level signal at all. Broadcast and cable exposure cannot be tied to individuals, which is why the legacy TV measurement industry built panels, and why panel limitations carry into CTV.
The result is that platform-reported ROAS for TV and CTV is not just noisy, it is structurally inflated. The only way to know what television actually contributes is to measure causally: what happens to your business when the ads run versus when they do not. For a deeper primer on why platform lift numbers fall short, see Why Platform Lift Tests Are Not Good Enough for Measuring Incrementality.
The Four Methodologies for Measuring TV and CTV Incrementality
Every approach to TV and CTV measurement falls into one of four families. They differ in what data they require, whether they establish causality, and how well they survive a privacy-first world.
1. Geo holdout and geo incrementality tests
A geo holdout test turns TV or CTV spend off in a scientifically selected set of test markets while the rest of the country continues business as usual. A counterfactual model, built from roughly two years of historical transaction data in the untested markets, predicts what conversions in the test markets would have been if the ads had kept running. The gap between predicted and observed conversions is the incremental impact of the channel.
Because geo tests read results from transaction data (your actual orders), they require no user-level tracking, no device graph, no cookies, and no cooperation from the ad platform. They work identically for linear TV, streaming, YouTube, and every walled garden. A scale test inverts the design, increasing spend in test markets to measure the return on incremental investment. For the full mechanics, see How to Run Geo Testing for Marketers: A Step-by-Step Guide.
2. Platform and publisher lift studies
Most major CTV sellers offer their own lift products, typically brand lift surveys or conversion lift studies run inside their walled garden. These are better than raw last-touch reporting, but they carry two structural problems. First, the platform controls the test design, the holdout construction, and the reporting, which creates an obvious incentive conflict. Second, the study only sees conversions the platform can match to its own identity graph, so the denominator is incomplete. Platform lift studies are useful as a directional signal and as a negotiating input, not as ground truth.
3. Panel and ACR-based attribution
TV-specialist measurement vendors, such as Tatari and iSpot, build measurement on viewership panels and automatic content recognition (ACR) data licensed from smart TV manufacturers. ACR identifies which households saw which ads, and the vendor matches exposed households against outcomes like site visits and purchases, often using baseline-versus-exposed comparisons to estimate lift.
This approach delivers fast, granular reads, down to the creative, daypart, and network level, which makes it genuinely useful for in-flight optimization of TV buys. Its limits are the limits of its inputs: ACR coverage is partial and skewed toward certain TV brands, household-to-outcome matching relies on the same identity graphs that privacy changes are eroding, and exposed-versus-baseline comparisons are correlational designs that can absorb selection bias (households that see more TV ads differ systematically from those that see fewer). Panel-based attribution estimates who converted after exposure. It does not establish what would have happened without the exposure.
4. Media mix modeling (MMM)
Marketing mix modeling estimates TV and CTV contribution statistically, using historical spend and outcome data. Modern MMM handles television reasonably well in one specific respect: adstocking, the transformation that spreads TV’s delayed impact across subsequent weeks, was invented for exactly this problem.
MMM’s weakness on TV is identifiability. TV spend is often lumpy, seasonal, and correlated with other upper-funnel investments, which makes it hard for a regression to separate TV’s effect from everything else that moved at the same time. An uncalibrated MMM can assign TV almost any coefficient the data will tolerate. The fix is experimental calibration: geo test results anchor the model’s TV estimate to ground truth, and the model then extends that truth across every week you do not have a live experiment. This is the logic of the incrementality, attribution, and MMM decision tree.
Comparison: TV and CTV Measurement Methodologies
| Geo holdout tests | Platform lift studies | Panel / ACR attribution | MMM | |
|---|---|---|---|---|
| Establishes causality | Yes, experimental | Partially, within the walled garden | No, correlational | Only when calibrated with experiments |
| Data required | Transaction data by geography | Platform-side exposure and conversion data | ACR / panel data plus identity matching | 2+ years of spend and outcome history |
| Privacy durability | High, no user-level tracking | Low to medium, depends on platform identity | Low to medium, depends on identity graphs | High, aggregate data only |
| Works across all TV sellers | Yes, channel-agnostic | No, one walled garden at a time | Mostly, where ACR coverage exists | Yes |
| Speed and granularity | Weeks per test, tactic-level | Weeks, campaign-level | Near real time, creative-level | Ongoing, channel and tactic-level |
| Independent of the ad seller | Yes | No | Yes | Yes |
| Best used for | Ground-truth incrementality and budget decisions | Directional validation | In-flight buy optimization | Always-on planning, calibrated by tests |
The takeaway is not that one methodology wins on every row. It is that only one row of capabilities produces a number a CFO should trust for budget decisions, and that row belongs to controlled experiments. Everything else is either optimization tooling or a model that needs experimental calibration.
How a Geo Holdout Test Works for TV and CTV
A well-designed TV or CTV geo test follows the same sequence regardless of whether you are testing linear, streaming, or YouTube.
- Market selection. The highest-volume markets, typically the top half of markets by historical conversion volume, are reserved as modeling markets used to build the prediction. From the remaining markets, an optimization algorithm selects the test cell that minimizes prediction error and bias across historical validation periods. Test markets usually represent somewhere between 7% and 40% of national sales. Random assignment does not work here: with only around 210 media markets in the US, random splits produce structurally imbalanced groups, which is why rigorous geo testing uses a synthetic control approach with optimized market selection. More on that design choice in Synthetic Controls Aren’t Enough: Why Geo Testing Requires Representative Market Selection.
- Treatment. For a holdout, TV or CTV spend is turned off in the test markets. CTV makes this operationally easier than linear ever was: most streaming platforms support geographic targeting exclusions, so the holdout can be implemented directly in the buying platform. For a scale test, spend is increased in the test markets instead.
- Duration. Tests typically run four to eight weeks. The duration must cover the full lag between exposure and conversion, which is longer for TV than for search or social, and it must be long enough for the design to detect the expected effect size. Upper-funnel channels with smaller expected lift require longer windows. See How Long Should an Incrementality Test Run for the underlying power math.
- Measurement. Weekly observed conversions in the test markets are compared to the model’s predicted conversions. For a holdout, incrementality is proven when predicted exceeds observed: turning the ads off created a measurable shortfall against the baseline, and that shortfall is the demand the channel was generating. The result is reported with a confidence interval and a statistical significance level, so you know how much weight the number can bear.
- Calibration. The test result becomes a calibration input for always-on measurement, adjusting platform-reported conversions to their true incremental value every week going forward, not just during the test window.
What a Real Streaming TV Holdout Looks Like
Here is an example from Measured’s testing environment that illustrates how the mechanics above translate into numbers a finance team can act on.
An enterprise ecommerce retailer ran a four-week geo holdout on its streaming TV spend. Streaming ads were paused in the test markets while the rest of the country continued business as usual.
- The counterfactual model predicted 117,208 orders in the test markets over the window. Observed orders came in at 114,906, a shortfall of 2,302 orders that represents the demand streaming TV had been generating.
- That shortfall translates to a 1.96% contribution to total orders (80% confidence interval of 1.75% to 2.18%), meaning streaming TV was responsible for roughly 2% of the business during the window.
- The test adjustment factor came in at approximately 52%: only about half of the conversions the streaming platform reported were actually incremental. The platform was over-crediting itself by nearly 2x.
- After applying that correction, streaming TV delivered a $2.46 incremental ROAS, with an 80% confidence interval of $2.20 to $2.73, at 90% statistical significance.
Two things make this example instructive. First, the platform over-reporting and the positive result coexist: streaming TV was genuinely working even though the platform’s own numbers exaggerated by how much. Marketers who dismissed the channel because they distrusted the reporting would have cut a profitable investment; marketers who took the platform numbers at face value would have over-invested. Second, none of these numbers required tracking a single user. The entire result was derived from geographic transaction data, which means it holds up regardless of what happens to cookies, device identifiers, or attribution windows.
Measured’s Q1 2026 benchmark data shows why results like this reward the brands that go looking for them. Among the 49 enterprise advertisers in the analysis (revenue above $200 million), median CTV incremental ROAS rises to $3.18, roughly 40% higher than YouTube and nearly double Linear TV, and those brands achieve it with a median spend only modestly above the overall group. The gap is not budget. It is allocation discipline built on incrementality measurement. The full breakdown, including the TV maturity curve that separates baseline video buyers from strategic CTV deployment, is in the downloadable report CTV is Mainstream. Optimization Just Hasn’t Caught Up. For a brand-side account of putting this into practice, see the Patagonia CTV incremental ROAS webinar.
From Single Tests to Always-on TV Measurement
A geo test tells you what TV contributed during the test window. It does not, by itself, tell you what TV is contributing this week. Conditions change: new creative, new streaming partners, different flighting, seasonal demand shifts.
The mature version of TV and CTV measurement combines the two strengths this guide has covered. Geo experiments provide periodic causal ground truth per channel and tactic. A media mix model, calibrated by those experiments, extends the ground truth across every week and every dollar, producing incrementality-adjusted ROAS continuously rather than only when a test is live. Each new test refreshes the calibration, so the model never drifts far from experimental reality.
This is how Measured approaches television for its customers: geo holdout and scale tests establish the causal contribution of each TV and CTV tactic, those results calibrate the Measured Incrementality Model (MIM), and the Cross-Channel Dashboard reports incrementality-adjusted performance daily alongside every other channel. The approach is validated across 25,000+ experiments and $35B+ in optimized ad spend for 160+ enterprise brands, and it is the reason Comcast’s Universal Ads named Measured its strategic incrementality measurement partner. For a brand-side proof point, see how Free Fly validated Universal Ads with Measured.
Frequently Asked Questions
What is the best way to measure CTV incrementality?
Geo holdout testing is the most rigorous way to measure CTV incrementality because it is a controlled experiment based on transaction data rather than user tracking. Spend is paused in scientifically selected test markets, and the gap between predicted and observed conversions reveals the channel’s true contribution. Panel-based attribution and platform lift studies are useful supplements for optimization, but only experiments establish causality independently of the ad seller.
How does holdout testing work for TV advertising?
In a TV holdout test, TV or streaming spend is turned off in a set of test markets while the rest of the country continues as usual. A model built on roughly two years of historical data predicts what conversions in the test markets would have been with the ads still running. If observed conversions fall measurably below the prediction, that shortfall is the incremental demand the TV spend was generating. Results are reported with a confidence interval and a statistical significance level.
Can you measure TV ad ROI without user-level tracking?
Yes. Geo-based incrementality testing measures TV impact entirely from aggregate transaction data by geography, with no cookies, device graphs, or identity resolution. This makes it fully privacy-durable: results are unaffected by attribution window changes, signal loss, or platform privacy policies, and the same method works identically across linear TV, streaming, and every walled garden.
What is the difference between panel-based TV attribution and geo testing?
Panel and ACR-based attribution (the approach used by TV-specialist vendors like Tatari and iSpot) matches households that saw TV ads against outcomes, producing fast, granular reads that are useful for optimizing buys in flight. But it is a correlational design that depends on partial ACR coverage and identity matching. Geo testing is an experimental design that measures what happens to actual transactions when ads are removed or scaled, which establishes causality and does not depend on tracking coverage. Many brands use both: panels for tactical optimization, geo experiments for budget-level truth.
How long should a CTV incrementality test run?
Most TV and CTV geo tests run four to eight weeks. The window must cover the full lag between ad exposure and conversion, which is longer for television than for lower-funnel channels, and it must be long enough to detect the expected effect size at the desired statistical confidence. Channels with smaller expected lift or noisier baselines need longer tests.
Which CTV platforms can be measured with incrementality testing?
All of them. Because geo testing reads results from your transaction data rather than platform-reported conversions, it works for YouTube, Hulu, Roku, Netflix, Amazon, Universal Ads, programmatic CTV through DSPs, and linear TV. The only operational requirement is the ability to exclude or scale spend by geography, which nearly every streaming platform supports.
