[{"data":1,"prerenderedAt":47},["ShallowReactive",2],{"article-en-holdout-and-geo-experiments":3},{"slug":4,"locale":5,"title":6,"description":7,"category":8,"categoryKey":8,"readMinutes":9,"publishedAt":10,"body":11,"sources":37,"seoTitle":6,"seoDescription":7},"holdout-and-geo-experiments","en","Holdout Experiments and Geo Experiments: When You Need a Control Group, Not Another Attribution Model","How randomized holdouts and geographic tests create counterfactuals, where they fail, and how to decide whether the business is test-ready.","Measurement",12,"2026-08-09",[12,17,22,27,32],{"h":13,"p":14,"bullets":16},"The control group is an investment in knowing",[15],"A holdout deliberately withholds the marketing treatment from a comparable group. That can feel uncomfortable because the business is choosing not to advertise to some eligible people. But the “lost” exposure buys information: it creates evidence about what would have happened without the campaign.",[],{"h":18,"p":19,"bullets":21},"User-level holdout vs geo experiment",[20],"A user-level randomized holdout is powerful when the platform or system can assign eligible people to treatment and control cleanly. A geo experiment uses non-overlapping geographic areas as the experimental units and changes spend or exposure by area. Geo designs are especially useful when user-level tracking is impossible, undesirable or fragmented across channels.",[],{"h":23,"p":24,"bullets":26},"The hidden enemies: spillover and imbalance",[25],"Experiments break when control users are indirectly treated, when customers move across regions, when test geographies are structurally different, or when promotions and sales operations change unevenly. Randomization helps, but design quality still matters. The business must protect the experiment from operational contamination.",[],{"h":28,"p":29,"bullets":31},"Power matters more than enthusiasm",[30],"A test can be conceptually perfect and statistically useless if the expected lift is too small relative to natural variation. Before launch, estimate baseline volume, variance, minimum detectable effect, duration and business cost. If the test cannot realistically detect the effect that matters, redesign it before spending the budget.",[],{"h":33,"p":34,"bullets":36},"Experiments should answer decisions",[35],"The best experiment ends with a decision rule: if incremental revenue exceeds a threshold, scale; if the confidence interval includes commercially unattractive outcomes, do not scale yet; if the result is inconclusive, decide whether more data is worth the opportunity cost. An experiment is not successful because it produced a p-value. It is successful because it reduced decision uncertainty.",[],[38,41,44],{"title":39,"url":40},"Google Research — Measuring Ad Effectiveness Using Geo Experiments","https:\u002F\u002Fresearch.google\u002Fpubs\u002Fmeasuring-ad-effectiveness-using-geo-experiments\u002F",{"title":42,"url":43},"Google Research — Incremental ROAS with Paired Geo Experiments","https:\u002F\u002Fresearch.google\u002Fpubs\u002Frobust-causal-inference-for-incremental-return-on-ad-spend-with-randomized-paired-geo-experiments\u002F",{"title":45,"url":46},"Google Ads Experiment Center","https:\u002F\u002Fsupport.google.com\u002Fgoogle-ads\u002Fanswer\u002F16856494?hl=en",1786584240476]