
Measurement Design
Incrementality and media experiments
Plan a media incrementality test around a clear decision, credible comparison, protected delivery and a result that includes uncertainty.
Incrementality asks what a media change caused compared with what would have happened without it. A campaign-period sales increase alone cannot answer that question. A credible experiment changes media delivery for an assigned group and compares its outcome with a suitable control.
Start with the decision
Define the change you might make after the result. A holdback can test new activity against no new activity. A go-dark test withdraws an existing activity in selected units. A heavier-spend test estimates the effect of adding spend; it does not measure the total value of the existing campaign.
Choose the business outcome and counting rules before launch. For example, confirmed orders may be assigned to an Australian market by customer address and counted over an agreed response period, with cancellations handled consistently. Keep platform delivery and attributed conversions for diagnosis, but distinguish them from the agreed outcome.
Choose a comparison you can maintain
The assigned unit might be an eligible platform user or a geographic area. For a geographic test, check whether media can be varied by area, outcomes can be located consistently, enough suitable areas exist, and exposure is likely to cross boundaries. A user holdout depends on the platform maintaining assignment and on which campaigns are included in the study.
Random assignment can strengthen comparability. When geographic units are few or markedly different, a matched-market design may be considered, but its causal interpretation depends on the matching and analysis assumptions. Use pre-test outcomes to assess fit and the effect the design could reasonably detect. Historical fit is a design check, not a lift result.
Treatment and control regions should be non-overlapping, with each region receiving its assigned condition through geo-targeted advertising.
Practical constraints may also rule randomisation out, such as running a smaller-scale experiment within a set budget or keeping particular geos in specific groups. A matched-market approach instead searches for group assignments that appear to satisfy the assumptions of a time-based regression model (Kerman et al., 2017); where those assumptions hold, the recommended designs lead to straightforward causal estimates.
Key Design Considerations for Geo-Experiments
- Non-overlapping regions
- Required to avoid spillover effects between treatment and control
- Pre-test fit assessment
- Use historical data to check similarity between treatment and control before launch
Specify the design before launch
The design workflow moves through preparing the pre-test data, setting the experiment parameters, evaluating whether a candidate design is viable, then visualising and exporting the chosen design.
Experiment design parameters are set before launch, and experiment type is recorded as holdback, go-dark or heavy-up. Cell count defaults to one treatment group paired with one control group; above one it becomes a multi-cell study, and each cell can carry a different type. Pairing go-dark in one cell with heavy-up in another is a common budget-neutral setup.
Duration typically covers at least one purchase cycle for the product and excludes any cooldown periods; a documented example is a 28-day heavy-up experiment. The analysis method is chosen for design selection — time-based regression, synthetic control or synthetic difference-in-differences — and locations are assigned by random or stratified sampling, with candidate designs ranked for the number you request.
Prepare the pre-test data
A prepared pre-test dataset is the prerequisite for designing the study. Required columns are a daily date, a location such as a market area, and conversions — usually raw conversion counts or revenue, typically from the advertiser's CRM. Spend is optional, but a multi-cell design needs a separate spend column per cell.
Protect the test and report its limits
Record the assigned units, dates, outcome, campaign changes and intended spend difference. Monitor whether delivery reaches controls or withheld spend shifts into other areas. Log material differences in promotions, prices, stock and other media rather than changing the analysis quietly. Allow the agreed response period to finish before treating a result as final.
Report the estimated additional outcome, the baseline used for relative lift and the method's uncertainty measure. If reporting incremental return on ad spend, define both incremental conversion value and incremental cost. An interval that includes no effect does not establish that the true effect is zero; the estimate may lack precision.
The final decision also depends on whether the planned intervention occurred and whether the tested areas, dates and spend level resemble the proposed rollout.
In this guide
- Choosing a suitable geography for a media testCheck targeting, outcome data, pre-test patterns, spillover and precision before choosing areas for a media experiment.
- Comparing a holdout with ordinary historical reportingSee what an ordinary historical media report can show, what a holdout adds and when neither supports a confident causal claim.
- Estimating spillover between exposed and control audiencesMeasure observable leakage in a media test, disclose unknown exposure and assess how spillover may affect the decision.
- Reporting uncertainty in an incremental lift resultPresent incremental lift with its outcome unit, baseline, uncertainty range, design limits and decision threshold.


