Geo holdout testing — Mara
| |

Geo Holdout Testing for Affiliate Programs: A Step-by-Step Runbook

Operators often wait for a perfect measurement stack while paying coupon partners full commission - week after week, the dashboard looks fine and the margin quietly bleeds. If you already know why last-click lies and why geo-holdout tests work for smaller programs, open your spreadsheet. This is the how. For the conceptual why - incrementality vs attribution, why regions instead of users - read the practical overview of geo holdout testing for small programs. That piece covers the foundation. This one is the runbook: pick regions, suppress cleanly, run long enough, read the right metrics, act on the result.

What This Guide Is and What You Need (Prerequisites)

This guide is for mid-sized affiliate programs: the ones big enough to have a coupon or cashback question, but nowhere near the scale for user-level holdouts or enterprise lift studies. The stack is refreshingly dull. You need access to your affiliate network’s geo or partner controls (Impact, CJ, or similar), Google Analytics or your order management system for regional revenue data, a spreadsheet, and the willingness to suppress a partner type for four to six weeks.

Geo-holdout testing is what you use when you’re already running a channel and want to see what happens if you stop. That framing comes from Funnel’s definition of geo-holdout tests. And because CJ says it plainly, “incrementality is not attribution,” keep one reminder in front of you: attribution tells you who touched last; incrementality asks whether the sale needed them at all. So what you actually need to begin:

  • Admin access in your affiliate network to apply geo restrictions or pause partner-level traffic by region.
  • GA or OMS reports segmented by state, DMA, or country, enough to pull weekly revenue and conversion rates by geography.
  • A blank spreadsheet and a change log template. Boring, required.

If you don’t have control over partner suppression by geo, this method won’t work and you’ll need a different approach. Now let’s build the test.


Before You Change Anything: Define the Hypothesis and the Unit of Test

Segment Partners by Influence Role, Not Network Label

Split partners by the job they actually do before you scope the test. Introduce partners create discovery at top of funnel: editorial reviews, niche content, creator placements. Persuade partners work mid-funnel through comparison posts, buying guides, and category roundups. Close partners sit at checkout: coupon, cashback, and extension players.

A geo holdout on “all affiliate” fails by design because close-role partners dominate last-click attribution and blur the read on introduces and persuaders. Scope the test to one role, or better, one named partner, before you touch any settings.

The question you’re answering isn’t “does affiliate work?” It’s “does this partner type actually create demand, or is it mostly capturing demand that was already heading to my checkout?”

Measurement Ladder: When to Stop Before a Geo Holdout

Not every program should jump straight to suppression. Run the cheaper layers first:

  • Layer 1: new-to-file rate KPI for acquisition vs recycled demand by partner. Directional only, no suppression required. If NTF is already high and stable, you may not need a holdout.
  • Layer 2: branded-search overlap, time-to-convert, prospecting vs retargeting vs branded-query patterns in paid search, and the five coupon cannibalization signals linked below. A coupon partner can look fine on prospecting while retargeting or brand-query overlap tells a different capture story. This catches checkout-layer partners without a test window. If those signals point to capture, move to Layer 3.
  • Layer 3: matched geo holdout. This is the causal read, and the rest of this guide is the runbook.

Stop at Layer 1 or 2 when the answer is already clear enough to act on. Commit to Layer 3 when you need causal evidence before changing commissions or partner terms.

Write the hypothesis in one sentence before you touch any settings:

If we suppress Partner Type X in Region A for six weeks, total revenue in A will fall less than affiliate-attributed revenue from X falls, relative to matched Region B where X runs normally.

That’s the test. Now pre-register your metrics. Baseline metrics to pre-register: total revenue (not just web, if you have retail or marketplace sales by geo, pull those too), orders, new-to-file customers if you can, and affiliate-attributed revenue for that partner type. Set a decision rule now. For example: if total revenue in the holdout region drops less than 5% while affiliate-attributed revenue from the partner drops by 30% or more, the partner is largely non-incremental. If total revenue drops materially, say, over 8-10%, and the control region stays flat, the partner is likely incremental. These percentages are illustrative decision heuristics for mid-sized programs, not universal statistical cutoffs, and should be calibrated to your program’s baseline volatility and risk tolerance.

Decide what “inconclusive” means and what you’ll do. Usually it’s extending the window, tightening the match, or narrowing to one partner. Write that down too.

Practically, announce the test internally before you start: finance, the marketing lead, anyone who might later ask why coupon-attributed revenue cratered in Texas. If you plan to adjust commissions afterward, decide the partner communication rules now. Trust breaks when you change terms silently after a test. This is the part that matters.

The partners most likely to fail an incrementality test are the ones showing checkout-layer patterns: low new-to-file rates, sub-10-minute conversion windows, high overlap with content traffic. The five signals that coupon partners are cannibalizing commissions are worth checking before you commit to a holdout. That way you’re not testing blind.

✓ Checkpoint: A single-sentence hypothesis, a chosen partner type, and a documented decision rule sit in your spreadsheet. No network settings touched yet.


Step 1: Select Matched Regions (Where Most Tests Quietly Fail)

Parallel trends check chart: control and test region revenue move together in the pre-period, then the test line drops after partner suppression while control stays flat.
If your pre-period lines don’t stick together like this, your match is a story, not a measurement. Fix the pair before you flip the switch.

This is where the whole thing usually falls over. Match on movement together, not identical levels. You need regions that trend similarly in total revenue, conversion rate, seasonality, and channel mix. Rockerbox recommends matched geographies that behave similarly on customer composition, channel performance, and overall conversion rate.

Practical unit: US state or DMA if you have that granularity; country if you’re global. Avoid regions with local events, weird promo calendars, or one-off launches. Those will turn your result into a story, not a measurement.

Now run a parallel-trends check. Plot weekly total revenue for your candidate test and control regions over the 4-8 weeks before the test start. If the lines diverge before you’ve done anything, one region flat, the other climbing, that pair is garbage. Remove it. Amsive’s guidance on defensible geo holdout tests is clear: parallel pre-treatment trends are the non-negotiable requirement for a causal read. Spreadsheet columns to start with: Region | Pre-period average weekly revenue | Trend slope | Conversion rate | Affiliate-attributed revenue share | Notes. You’ll add historical data and the parallel-trends chart.

Pitfalls that kill the read before it starts:

  • Survivorship bias: keeping only the “pretty” matches after the fact. Pick the pair before you see test data.
  • Averaging across too many regions until broken pairs disappear.
  • Mixing a high-intent core market with a casual or seasonal market in the same pair. Revenue levels can look similar while movement diverges every promo week.
  • Using one iconic city, “we’ll just use Chicago,” as the national control. A single city is rarely representative.
  • Trusting an automated matched-market score without opening the pre-period chart. Vendor matching is a starting filter, not the parallel-trends proof.
  • Forgetting to avoid matched-market pitfalls like unobserved demographic drift, local competitor store openings, and mismatched offline media weight.

Reject a pair if:

  • The pre-period trends diverge by more than a stable band. Eyeballing a 5-7% gap is enough to be nervous. These thresholds are rules of thumb for mid-sized programs; larger programs may use tighter bands, smaller ones wider.
  • One region has a known local event or store opening during the test window.
  • Conversion rate differs by >20% consistently in the pre-period. Again, a heuristic, not a hard cutoff.
  • The affiliate partner type you’re testing already has heavily skewed regional adoption, such as a cashback portal used mostly on one coast.

✓ Checkpoint: You have a documented matched pair with a parallel-trends plot saved, and any discarded pair has a written reason why.


Step 2: Set Up Suppression Cleanly (and Prove Control Is Actually Dark)

Goal: make the test region truly dark for the partner type, without contaminating the control or letting other channels pick up the slack in ways you can’t measure. Evaluating cashback or coupon suppression requires new-to-file plus branded-search overlap plus time-to-convert in the same window, not affiliate-attributed rows alone. Document every single change: who, when, which links or placements, which network UI toggles.

For coupon, cashback, or extension partners: geo-restrict their tracking links or remove the test region from approved traffic. Many networks (Impact, CJ, Awin) offer partner-level traffic or geo restriction controls. Check your network’s documentation for the exact toggle; it usually sits under partner settings, traffic restrictions, or geo targeting. Where your network supports partner-level geo or traffic restrictions, you block traffic from specific states or countries. If that’s not available, you may need to create a separate tracking link that only serves outside the holdout region.

For content partners: pause placements or creative only in the holdout region if you can. If that’s not possible, exclude that partner type from this test design. Don’t run a test blurred across partner types; you’ll lose the causal read.

Now the contamination audit. It’s boring but mandatory. Verify that the control region stays dark: no leftover deep links pointing to the affiliate code, no email sends with the affiliate URL, no retargeting audiences that reintroduce the partner’s cookie, no DSP/affiliate bleed. One leaked link and your six-week holdout becomes an expensive diary entry. A common contamination pattern: an email team sends a promotional blast using a coupon code that contains the partner’s tracking parameter, inadvertently reintroducing the affiliate cookie into the holdout region.

Start a change log now:

Date Change Region Partner Expected Effect Owner
YYYY-MM-DD Geo-restricted links for Partner X Holdout (TX) Coupon X Affiliate clicks from TX stop You
YYYY-MM-DD Verified GA traffic for TX shows no Partner X referrer Holdout (TX) Coupon X Confirm suppression You

When you’re testing coupon-like partners, you’re often investigating the same checkout-layer interception patterns that surfaced in the Honey lawsuits: the extension fires at checkout, claims credit, and the original content partner earns nothing. The five operational changes after Honey-style attribution hijacks include tightening tracking windows and auditing which partners actually added value. A geo holdout is the data version of the same instinct.

✓ Checkpoint: The partner is suppressed in the test region. Your change log is dated and owned. The control region is verified dark, and you’ve spot-checked that no affiliate codes are leaking.


Step 3: Run Long Enough and Log Confounders Weekly

Minimum four weeks; six is better. Short windows get pushed around by normal variance and conversion lag. Funnel’s geo-holdout article uses a six-week test window for its fictional example, and Rockerbox similarly stresses that tests need enough time to measure latent impact. Price-sensitive shoppers that use coupons might convert within hours, but a cashback user might wait for a statement cycle, and your site might have a 14-day return window. The test needs to capture the full purchase and reversal cycle.

Open a weekly confounder log. Boring, required. Every week, record:

  • Paid media changes by region, especially search and social budget shifts.
  • Email and SMS sends, by region if you segment.
  • Sitewide promotions or discount changes.
  • Price changes or stock-outs on key SKUs.
  • Competitor activity that you hear about: a big sale, a launch.
  • Any site outages or checkout friction.

Price and stock-outs are especially dangerous. If your holdout region runs out of a top seller in week three, total revenue will drop for reasons that have nothing to do with the partner you suppressed. That’s a confounder, not a failure of the partner. Note it and consider extending the test.

Do not peek daily and “optimize” mid-test. If you pre-registered an interim look, fine. Otherwise, daily peeking leads to tweaking something: pausing the test early, adjusting budget in the control region. That contaminates the design. Let the test run.

✓ Checkpoint: The test is running with a scheduled end date. There’s a weekly log with confounders, and you haven’t touched the partner settings since launch.


Step 4: Read the Results the Right Way (Lift Math Without Theater)

Difference-in-differences readout table: compare test and control total revenue change, then estimate lift as delta test minus delta control, reading demand capture vs incremental.
Attributed drop is not the answer. Estimated lift is Δ_test − Δ_control on total revenue — then classify demand capture vs incremental.

When the test ends, your primary contrast is not the drop in affiliate-attributed revenue. It’s the change in total revenue, orders, or new customers in the holdout region versus the control. Worth repeating: if affiliate-attributed revenue for the suppressed partner collapses but total revenue stays mostly flat, the partner was capturing existing demand. If total revenue drops materially with the partner gone, the partner was likely incremental.

Apply difference-in-differences thinking: calculate the pre-period versus post-period change in total revenue for both regions. Then subtract the control region’s change from the holdout region’s change. That difference is your estimated lift, or loss, from the partner. No stats PhD required. For example, if the holdout region’s total revenue fell 2% post-treatment while the control region rose 1%, the net difference is a 3% drop attributable to removing the partner. If affiliate-attributed revenue from that partner was 10% of total revenue before, and it disappeared entirely in the holdout, then you were paying for a 10% share of revenue to generate only 3% incremental lift. That’s directionally clear: mostly demand capture.

You can confirm with iROAS: incremental revenue over incremental affiliate cost. If you spent $0 on commissions in the holdout region but lost just a bit of revenue, you might have an extremely high iROAS, which sounds great until you realize it means the partner was adding almost nothing.

Small programs get directional evidence, not lab-grade p-values. A 3-5% difference is noise. Treat a delta of 10% or more as the threshold for confidence. If the delta is in single digits and your order volume is under a thousand per region, call it inconclusive.

Measured’s geo-holdout for a premium home goods brand found Google Brand Search was 94% overvalued, with only 6% of sales actually incremental. That’s Google Search, not affiliate, but the structural lesson is the same: platforms routinely overstate causal value. Your holdout is the reality check.

If the result is genuinely inconclusive, go back to the measurement ladder first. Use Layer 1 and Layer 2 diagnostics before extending the window, tightening the regional match, or widening the test. If those still point to a capture problem but no clean read, extend the window, tighten the match, or narrow to one partner. If you’re stuck between a geo holdout and a multi-touch attribution dashboard you never use, remember: for most mid-sized programs, doing one clean geo test on one partner type beats six months of staring at MTA dashboards with volume too low to trust.

✓ Checkpoint: You’ve calculated a DiD-based revenue impact and can classify the partner as likely incremental, likely non-incremental, or too noisy to call.


Step 5: Act on the Result (or You Wasted Six Weeks)

Step 5 decision tree after a geo-holdout DiD readout: likely incremental, likely non-incremental, or mixed/noisy paths, plus always-do partner communication and schedule a retest.
Classify the signal, assign an owner, put a retest on the calendar — or you spent six weeks proving nothing operational.

Decision tree:

  • Likely incremental: Keep running. Grow if the economics justify. Now you have evidence to defend the spend.
  • Likely non-incremental: Narrow attribution windows, split commissions by customer type, reduce the base rate, or tier that partner type. Don’t cut outright without communicating, but align the payout with actual added value.
  • Mixed or noisy: Retest one partner. Don’t rewrite the whole program based on an ambiguous signal.

After the finding, the communication sequence is lead time, method, data, then terms. When you present to finance, show the matched regions, the parallel-trends plot, and the confounder log. When you talk to the partner whose commissions you’re adjusting, give lead time, show the method and the readout, honor any existing committed periods, then adjust terms. Partners who are actually incremental will have their own data. Partners who aren’t usually know what they are.

After one successful manual test, the graduate move is a rotating or always-on small dark geo: keep a low-traffic region suppressed and watch the signal continuously. That is a permanent monitoring design, not a starting point.

Incrementality insights expire. Attribution drift, new product launches, and shifting partner behavior mean a result from March might not hold in September. Plan a lighter re-check quarterly or after major program changes. If you get comfortable with the method, you can eventually set up a rotating small dark geo as a permanent low-cost signal.

✓ Checkpoint: An action is assigned, a partner communication plan exists, and a retest date is on the calendar.


Troubleshooting Common Issues

Broken parallel trends. Your two regions diverged in the pre-period. Go back to matching and pick a different pair. If you don’t have a good pair, you can still run a synthetic control, such as Google’s CausalImpact, but that’s a complexity jump.

Contamination. The holdout region is getting affiliate traffic despite your suppression: a deep link, an old email, a retargeting pool. Kill everything you can find and extend the test to wash out the residual.

Local promotion confounded the result. A competitor opened a store in your control region during week three. That shows up in the confounder log. If the effect is large, the test is invalidated and you need a clean window.

Duration too short. If you ran four weeks and saw a 4% delta with high variance, run six or eight. Small programs often need longer to smooth week-to-week noise.

What geo holdouts cannot do: they won’t give user-level precision, they can’t prove creative-level truths, and they may understate impact if there’s cross-region spillover. When geography cannot be split cleanly, a time-based on/off interrupter in the same region (partner dark for a window, then back on) is a weaker fallback. It confounds seasonality and promos, so treat that read as directional only, not a substitute for a matched pair. When you need precise user-level incrementality, escalate to user-level holdouts or marketing mix modeling. That’s a different resource class. Most “failed” geo tests weren’t failures; they were rushed designs pretending to be experiments.


Downloadable Checklist

Tape this to the monitor. If every box isn’t checked, you’re not testing, you’re guessing with geography.

  • ☐ Hypothesis and partner scope written down
  • ☐ Matched pair selected, parallel-trends plot saved
  • ☐ Suppression method documented, control region verified dark
  • ☐ Pre-period baseline data exported (4-8 weeks of weekly revenue, CVR, orders)
  • ☐ Weekly confounder log started
  • ☐ Test duration committed (4-6 weeks, end date set)
  • ☐ Primary metrics and decision rule pre-registered
  • ☐ Post-test DiD-style readout scheduled on the calendar
  • ☐ Action owner assigned for the decision
  • ☐ Retest date on calendar (quarterly or after major program shifts)

And remember: geo holdouts show directional causal evidence, not user-level precision. They can’t account for cross-region spillover or prove why one creative works. But for the question “is this partner type adding demand or just intercepting it?” they’re the most honest answer a mid-sized program can get. Don’t be the operator who runs the test, gets a clear answer, and changes nothing.

Affiliate Intelligence

Get the next actionable tactic by email

One practical affiliate marketing idea per week. No filler. No spam.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *