Introduction
Ask three vendors how long your incrementality test should run and you will get three confident answers. Ask a statistician and you will get a question back: how much signal are you trying to detect, and how much noise is it hiding in? That question is the actual answer. Duration is not something you pick; it is something you calculate, and every week you get wrong in either direction has a cost. Too short, and you spend real money on a test that ends in a shrug. Too long, and you burn held-out revenue buying precision you did not need.
This guide walks through the calculation in plain terms, gives directional planning tables, and covers the two execution mistakes that quietly invalidate more tests than bad math ever does.
The Three Quantities That Determine Duration
Every test duration question reduces to a relationship between three quantities:
- Baseline volume: how many conversions flow through the test footprint per week. Statistical noise shrinks as volume grows. A brand pushing 20,000 weekly orders through its test markets accumulates evidence roughly four times faster than a brand pushing 1,250, because uncertainty scales with the square root of sample size. Quadrupling volume halves the noise. This is why the same tactic can be testable in four weeks for a national retailer and untestable in a quarter for a niche brand.
- Target lift and minimum detectable effect (MDE): what matters versus what the test can see. These are two different quantities, and confusing them is the planning mistake teams make most often. Your target lift is a business input: the smallest lift that would change your decision, derived from economics. If a channel needs to drive at least 8% lift to clear breakeven iROAS at its current spend, 8% is your target lift. Your MDE is a statistical output: given a specific design (sample, variance, confidence, power), it is the smallest true lift that test could reliably detect. The planning rule connects them: design the test so its MDE is at or below your target lift. A design whose MDE is 15% will call a genuinely profitable 10% lift “not significant,” and you will cut a working channel. Teams go wrong by anchoring the design to the lift they hope for rather than the breakeven lift that matters, which produces tests that can only see home runs. Pushing the MDE down is expensive: halving it roughly quadruples the sample you need, which is why chasing precision on small effects escalates fast.
- Confidence and power: how sure you need to be. Convention is 90 to 95% confidence (protection against declaring a lift that is not real) and 80% power (probability of detecting a lift that is real). These are dials, not laws. A reversible creative-budget decision can tolerate 85% confidence; a decision to cut a seven-figure channel should not.
Volume, target lift, and confidence jointly determine required sample size: you size the test so its MDE meets the target. Divide required sample by weekly volume, and duration falls out.
A Worked Example
A DTC apparel brand wants to test whether its Meta prospecting spend is incremental.
Setup:
- Test footprint: markets producing about 5,000 orders per week
- Economics: at current spend, the channel breaks even at roughly 10% lift, so that is the target lift, and the test must be designed with an MDE at or below 10%
- Standards: 90% confidence, 80% power
A power analysis on this setup (accounting for the week-to-week variance in the brand’s demand) lands at roughly six weeks of runtime. Now watch how the answer moves when the inputs move:
- If the brand only cared about detecting 20%+ effects, the same footprint could read in about two to three weeks (a bigger effect needs far less data)
- If weekly volume were 1,250 instead of 5,000, detecting that same 10% lift would take roughly four times as long, pushing past 20 weeks, at which point the right move is a bigger footprint, not a longer calendar
- If demand is highly volatile (heavy promo cycles, seasonal swings), variance rises and every cell in this math gets longer
The numbers above are illustrative; the structure is not. In practice this calculation is run through simulation against your actual market-level history, which is exactly what simulation-based market selection does before a dollar is held out.
Directional Planning Table
For early planning conversations, before a proper simulation, these ranges are a reasonable starting map. Assumes well-matched markets, stable demand, 90% confidence, 80% power:
| Weekly conversions in test footprint | Detecting ~20% lift | Detecting ~10% lift | Detecting ~5% lift |
|---|---|---|---|
| 20,000+ | ~2-3 weeks* | ~3-4 weeks | ~6-8 weeks |
| 5,000 | ~3-4 weeks | ~5-7 weeks | ~10-14 weeks |
| 1,000 | ~5-8 weeks | ~12+ weeks | Redesign the test |
*Subject to the floors below; almost no test should run under four weeks regardless of what the power math allows.
Treat this as a compass, not a contract. Your variance, seasonality, and market structure move every cell. The table’s real job is to show the shape of the tradeoffs: effect size dominates, volume buys speed, and small-effect hunting on thin volume is a redesign problem, not a patience problem.
Why There is a Floor Around Four Weeks
Even when volume is high enough that the statistics clear in two weeks, two practical floors apply:
Conversion lag. Media influences purchases that land days or weeks later. End the measurement window too early and the lagged conversions the media caused are missing from the read, which systematically understates lift, and it understates it most for exactly the upper-funnel tactics that are hardest to measure any other way. Rough lag expectations by category: impulse and low-ticket ecommerce, a few days; considered apparel and beauty, one to two weeks; furniture, mattresses, and high-ticket home, two to six weeks; B2B and long-cycle considered purchases, long enough that geo testing may need cohort-based outcomes rather than immediate revenue.
Whole demand cycles. A test should span complete weeks (weekday and weekend patterns weigh equally in test and control) and, ideally, avoid straddling a demand regime change like the start of a major promo period. A test that runs two and a half weeks, or begins the Tuesday before Black Friday, has a built-in tilt no statistics can remove. For the seasonal question specifically, see incrementality testing in Q4.
The Early-Stopping Trap
The most common way a well-designed test gets ruined has nothing to do with design. It is a dashboard habit: watching the result daily and calling the test the moment it crosses significance.
Here is the problem. A fixed-duration test’s false-positive guarantee assumes you look once, at the end. Look every day, and you give random noise dozens of chances to wander across the significance threshold, and noise takes those chances. Check daily for six weeks and your real false-positive rate can quietly triple or worse. Teams then scale budget on “confirmed” effects that were never there, which is how bad experiments lead you the wrong way while feeling rigorous the whole time.
The discipline: size it, launch it, read it once at the planned end. If your organization genuinely needs valid interim reads, that exists, but it requires sequential testing methods designed for repeated looks, with wider thresholds at each peek, agreed before launch. Informal peeking is not a lighter version of that; it is the failure mode.
The mirror-image error deserves a mention: extending a test that has not “hit significance yet” in hopes it will. That is the same sin in reverse, shopping for the answer rather than accepting the planned read.
When the Math Says "Too Long"
Sometimes the honest calculation returns 16 weeks and the business will not wait. In order of preference:
- Widen the footprint. More or larger test markets is the cleanest fix; it buys volume without touching rigor. See how many markets you need for a geo test.
- Test a louder treatment. A bold spend change (going fully dark in holdout markets, or a 50% budget step instead of 15%) creates a bigger effect that is faster to detect. Timid treatments are the top cause of underpowered tests.
- Test a bigger unit. Measure the channel, not the campaign; decompose later with follow-up tests once the big question is settled.
- Trade precision knowingly. A four-week directional read with a wide confidence interval can be the right business call, as long as the interval ships with the result and nobody presents the point estimate as truth.
- Question the method. If volume is thin enough that no reasonable design gets there, the channel may not be ready for a clean experiment yet; a triangulated approach that leans on MMM until scale arrives is more honest than a doomed test.
Pre-Launch Checklist
- Write down the decision: what result makes you scale, cut, or hold, and at what threshold. That threshold is your target lift, and the design must deliver an MDE at or below it.
- Pull weekly conversion volume for the candidate footprint and check its variance.
- Run the power analysis (or have your measurement partner simulate it) at your confidence and power targets.
- Add lag buffer for your category and round to whole weeks.
- If the answer is impractical, apply the levers above now, never mid-flight.
- Lock duration, launch, and read once.
Teams that follow this sequence get something more valuable than a number: they get a result the CFO cannot poke a hole in. For the execution side of that credibility, see the full guide to running geo tests step by step.
Key Takeaways
- Four to eight weeks is the common range, but duration is calculated from volume, minimum detectable effect, and confidence, not picked from a calendar
- Set your target lift from economics (typically breakeven lift, not the lift you hope for), then design the test so its MDE is at or below that target
- Halving the MDE roughly quadruples required sample; quadrupling volume halves the noise
- Nearly every test has a floor around four weeks, driven by conversion lag and whole demand cycles rather than statistics
- Never stop a fixed-duration test at the first significant reading; daily peeking multiplies false positives and scales budgets on phantom effects
- If the math says the test is too long, widen the footprint or bolden the treatment before you sacrifice rigor
Frequently Asked Questions
How long does an incrementality test take? Typically four to eight weeks. The exact duration is set by a power calculation: how many conversions your test footprint produces weekly, the smallest lift that would change your decision, and the confidence level you require. High volume and large expected effects shorten tests; thin volume and small effects lengthen them or require a larger footprint.
What is a minimum detectable effect (MDE)? The smallest true lift a given test design can reliably detect, determined by its sample size, outcome variance, confidence level, and power. It is a statistical property of the design, distinct from your target lift, the business threshold that would change your decision. Sound planning designs the test so the MDE is at or below the target lift; a test whose MDE sits above breakeven will report a real but profitable lift as “not significant” and lead you to cut a working channel.
What sample size do I need for an incrementality test? Enough conversions that the design’s MDE reaches your target lift at your chosen confidence and power, which depends on volume, effect size, and demand variance. As rough intuition: detecting a 10% lift takes roughly four times the sample of detecting a 20% lift. The precise number comes from simulation against your own market history, not a generic formula.
Can I stop an incrementality test early if results look significant? Not if the test was designed for a fixed duration. Repeatedly checking and stopping at the first significant reading inflates the false-positive rate severely, because noise eventually crosses any threshold. Read the test once at its planned end, or pre-commit to a formal sequential design built for interim looks.
Why do incrementality tests need at least four weeks? Two practical floors: conversion lag (media-driven purchases land days or weeks after exposure, and cutting the window early drops them from the read) and demand cycles (tests should cover whole weeks so weekday and weekend patterns balance across test and control). Both apply even when raw statistical power would clear sooner.
What if my business is too small for a readable test? Widen the test footprint, make the spend change bolder, or test at the channel level instead of the campaign level. If volume still cannot support a clean read, use an MMM-led or triangulated approach until scale supports experimentation; an underpowered test that finds nothing proves nothing.
