Geo Holdout and Incrementality Testing with Power Analysis

Cheveley PartnersProNew0 orders on this service
Analytics, Tracking and Attribution · Incrementality and geo-lift testing

Power analysis comes first: if the test cannot detect the effect you expect, we say so before it runs rather than after it returns nothing.

About this service

About a third of the geo tests we are asked to review could never have detected the effect they were commissioned to find. The arithmetic is unforgiving. With twelve treatment markets, a pre-period fit of the quality most accounts actually have, and four weeks of runtime, the minimum detectable effect usually lands somewhere between 7 and 12 percent. If the honest expectation for that channel is a 3 percent lift, the test will return no significant effect whether the channel works or not, and someone will cut the budget on the strength of it. So the first deliverable of every testing engagement here is a power analysis, and sometimes the answer is that the test should not be run. How we design: We take 18 to 24 months of daily conversions by geography, build the market pool, and run placebo tests across historical windows — assigning treatment to markets where nothing happened and measuring how often the method finds an effect anyway. That distribution is the real error bar, not the p-value the package prints. We use augmented synthetic control through Meta's GeoLift, and Bayesian structural time series through CausalImpact where the data structure suits it better. When the two disagree we report both and explain the disagreement rather than picking the friendlier one. The Gulf is the hard case, and it takes most of our design time. The UAE offers emirate-level targeting granularity with effectively two markets carrying the volume, so a UAE-only geo test is not a test; it is a before-and-after chart with a confidence interval bolted on. We will not sell one. Saudi Arabia is genuinely testable — Riyadh, Jeddah, the Eastern Province cluster, Mecca, Medina, Abha and Tabuk give enough separable units to build a synthetic control with a defensible pre-period fit. Kuwait, Qatar and Bahrain are single-market countries and have to be pooled into a regional design or measured another way entirely. When geography cannot carry the test we change method rather than lower the standard. Time-based switchback designs, alternating on and off in balanced blocks, work for channels with short conversion lags. Platform-side holdouts with public-service-announcement or ghost-ad controls work where the platform implements them properly. User-level holdouts through your own suppression lists work for CRM and retention spend, which is where the most overstated numbers in this industry live. Reading the result: You get the effect estimate with its interval, the incremental cost per acquisition it implies, the placebo distribution the estimate sits against, and a plain statement of what the test does not tell you. A geo test measures the incremental effect of the spend level you tested, in the markets you tested, in the weeks it ran. It does not license extrapolation to double the budget, and that sentence goes on the front page every time. What we refuse: We do not run tests shorter than the conversion lag plus two weeks; for fintech account approval and for sportsbook first deposits that means six weeks as a floor. We do not stop a test early because the numbers look good, and the stopping rule is agreed in writing before launch. We do not report a lift figure without the placebo distribution behind it. We do not run a test whose result you have already told us will not change a budget. Who this is not for: Advertisers confined to a single city or a single small country with no regional pooling option. Teams who need a result this month. Anyone who wants a test to settle an internal argument rather than change a spend decision — it will settle nothing, because whoever loses will attack the design, and they will usually have a point.

Scope

Target market
Worldwide, UAE and GCC, Saudi Arabia
Working language
English
Industry
Fintech, iGaming, Pets
Engagement model
One-off project
Turnaround
1 month or more
Seller type
Boutique agency

What the seller needs from you

  1. 1Can you export 18 to 24 months of daily conversions by city or region?
  2. 2Which channel or campaign do you want to test, and what do you expect it to deliver?
  3. 3Can you hold a spend change stable for six weeks without exception?
  4. 4Which markets are in scope, and what is the volume split between them?
  5. 5What decision will the result change, and who makes it?

Asked at checkout. Delivery time starts once you answer, not when you pay.

Reviews

No reviews on this service yet.

Reviews appear only after an order completes, and both sides review each other. Nothing here is seeded or bought.

Other sellers offering incrementality and geo-lift testing

See all →

Starting at $8,000