All insights
XINDAR INSIGHT

Did GEO Cause the Change? Designing a Test with Controls

To test whether a GEO change caused an outcome, define the treatment and outcome before editing, assign comparable pages or query clusters to treatment and control groups, collect repeated observations before and after release, and estimate the change in the treated group relative to the control. Record platform changes, SEO changes, seasonality, and spillovers.

Direct answer: To test whether a GEO change caused an outcome, define the treatment and outcome before editing, assign comparable pages or query clusters to treatment and control groups, collect repeated observations before and after release, and estimate the change in the treated group relative to the control. Record platform changes, SEO changes, seasonality, and spillovers. Without a credible counterfactual, report an association rather than uplift.

“Citations increased after we rewrote the page” describes a sequence. It does not identify a cause.

The answer engine may have changed its model, retrieval index, citation interface, or query routing. A competitor may have removed a page. Demand may have shifted. Traditional search ranking may have improved. The measurement vendor may have changed its prompt set. A result observed after treatment contains the treatment effect plus every concurrent change that reached the outcome.

A controlled design tries to estimate the missing quantity: what would have happened to the treated pages during the same period if the GEO change had not been released?

Begin with a causal estimand

A hypothesis is too vague if it says “improve AI visibility.” Specify the estimand: the exact effect you want to estimate.

Example:

Among 40 eligible English product-guide pages and a frozen panel of 120 buyer questions, what is the eight-week average effect of replacing unqualified specification prose with governed fact blocks on the probability that the page is cited in Microsoft Copilot, relative to unchanged matched pages?

This sentence defines:

  • units: product-guide pages and associated questions;
  • treatment: one named editorial change;
  • outcome: citation presence, coded per observation;
  • platform: one declared product;
  • population: 40 eligible English pages;
  • window: eight weeks;
  • counterfactual: matched pages left unchanged.

Change any of these after seeing results and the experiment answers a different question.

Choose an experimental unit that can receive treatment

GEO interventions can be assigned at several levels.

UnitSuitable treatmentMain risk
Pagefact-block rewrite, heading change, evidence sectionother pages may compete for the same query
Topic clustercoordinated page and internal-link updatefewer independent units
Market-language versionlocalized entity and evidence recordregional demand and platform availability differ
Prompt clusteranswer-intent coverage programtreatment may touch several pages
Time periodsitewide template or technical releaseplatform and seasonal changes are confounded

The assignment unit and analysis unit must agree. If a whole topic cluster is changed together, counting each page as an independent treatment exaggerates the effective sample size. If the sitewide template changes, there may be no untreated page left inside the site.

Randomized holdouts provide the cleanest comparison

When enough comparable units exist, randomly assign them to treatment and control before editing. Stratify first when important baseline differences exist. For example, pair pages by topic class, search demand, baseline citation rate, freshness, and conventional ranking, then randomize one page in each pair.

The control group is not neglected content. It is the estimate of background change. Both groups should receive routine maintenance required for accuracy and security. The treatment group receives the additional intervention being tested.

Do not select treatment pages because they recently fell. That creates regression to the mean: unusually poor observations often improve on their own. Do not place the strongest pages in treatment and the weakest in control. Random assignment protects against those choices, although small samples can still be imbalanced.

Pre-register decisions that are easy to change later

Write the analysis plan before release. At minimum, freeze:

  1. primary hypothesis and primary outcome;
  2. unit of assignment and eligibility rules;
  3. treatment specification and release checklist;
  4. control condition;
  5. platforms, regions, languages, prompts, and repeat count;
  6. baseline and follow-up windows;
  7. missing-answer, no-citation, timeout, and blocked-page rules;
  8. exclusion rules and quality failures;
  9. statistical model and uncertainty interval;
  10. stopping rule and minimum duration;
  11. secondary outcomes and multiple-comparison treatment;
  12. conditions that invalidate the experiment.

Pre-registration does not make a weak design strong. It prevents a team from trying many outcomes and presenting the most flattering one as if it had always been primary.

Measure repeated outcomes, not one answer screenshot

AI answers can vary across time, sessions, products, model versions, account context, and geography. A single answer is one draw from that changing process.

Use a frozen prompt panel with stable IDs. Run prompts on a fixed cadence and preserve the exact answer, citations, visible product or model label, date, region, language, and account condition. If a platform allows a reproducibility parameter or API seed, record it; do not assume it eliminates system variation.

Useful primary outcomes include:

  • citation presence for the treated URL;
  • source-selection rate across eligible observations;
  • supported-claim rate when citation quality is the objective;
  • correct entity-description rate;
  • explicit recommendation rate under a fixed coding rule.

Citation count per answer may be skewed and difficult to interpret. A binary event per prompt-run is often easier to audit. Whatever the outcome, preserve zeros: no answer, no retrieval, and no citation are distinct states and should not disappear from the denominator.

Estimate a difference in differences

A simple controlled pre/post estimate is:

Estimated effect = (treated after − treated before) − (control after − control before).

Suppose citation presence rises from 20% to 32% for treated pages and from 18% to 27% for controls. The treated increase is 12 percentage points, but the control increase is 9. The difference-in-differences estimate is 3 percentage points. Reporting 12 points as “GEO uplift” would attribute the background rise to the treatment.

This method depends on assumptions. Callaway and Sant’Anna’s work on difference in differences formalizes estimation with multiple periods and staggered treatment timing. A key identifying idea is parallel trends: absent treatment, the treated and comparison groups would have followed comparable outcome paths, potentially after conditioning on observed covariates.

For a GEO study, inspect pre-treatment trends. If treated pages were already accelerating while controls were flat, the control may not represent the counterfactual. Similar pre-trends do not prove the assumption, but visibly different pre-trends weaken it.

Staggered rollout can create useful controls

Operational teams often cannot withhold a change permanently. A randomized staggered rollout can preserve learning:

  • Wave A receives the change in week 1.
  • Wave B remains untreated until week 5.
  • Wave C remains untreated until week 9.

Not-yet-treated units can serve as comparisons for earlier waves. Treatment timing must be assigned before outcomes are observed. The analysis should account for treatment-effect heterogeneity across groups and time; a simple two-way fixed-effect average can be misleading when effects differ by cohort.

Anticipation also matters. If editors improve control pages while preparing their later rollout, those units are no longer untreated. The difference-in-differences paper explicitly discusses limited treatment anticipation. Record content work beginning before the public release.

When randomization is unavailable

A sitewide release may leave no internal holdout. You can still build a comparison, but the causal claim becomes more assumption-dependent.

Matched external or internal controls

Use unaffected markets, languages, content types, or topic clusters with similar pre-period behavior. Explain why the intervention should not affect them and why the same external shocks should.

Interrupted time series

Model the pre-release level and trend, then test whether the post-release path departs from the forecast. A short baseline or coincident platform update makes this fragile.

Synthetic control or structural time series

The CausalImpact paper proposes a state-space model that predicts the counterfactual response using pre-intervention data and contemporaneous covariates. Its authors discuss use when randomized experiments are unavailable and also discuss limitations. For GEO, candidate covariates could include untreated topic groups, conventional search exposure, page availability, and category-level citation activity—provided those series are not themselves affected by treatment.

These approaches estimate a counterfactual; they do not observe it. Results should state the model, controls, pre-period fit, assumptions, uncertainty, and sensitivity checks.

Protect the control from interference

One unit’s treatment can change another unit’s outcome. This is interference.

If two pages compete for the same grounding phrase, improving one may reduce citations to the other. New internal links may alter retrieval for both groups. A sitewide schema change can expose control pages. A press campaign can raise entity recognition across every page. In those cases, the experiment estimates a mixture of direct effects and spillovers.

Reduce interference by assigning whole topic clusters, separating markets, or measuring cross-unit effects explicitly. Keep a change log for:

  • redirects, canonicals, indexation, and robots controls;
  • internal linking and navigation;
  • structured-data templates;
  • major site releases;
  • external coverage and distributor syndication;
  • paid campaigns and demand shocks;
  • known platform or measurement changes.

Test one treatment package you can describe

“We did GEO” is not reproducible. Define the intervention as a checklist.

A fact-block treatment might require:

  • a 50–80 word direct answer;
  • explicit entity, attribute, value, unit, condition, and date;
  • a source linked beside the claim;
  • one governed comparison table;
  • corrected heading hierarchy;
  • unchanged title, URL, internal links, and schema during the primary test.

If the treatment changes content, title, URL, internal links, structured data, and promotion at once, the package may be worth testing, but the study cannot identify which component caused the result.

Use research benchmarks as priors, not promised effects

External GEO studies show why local controls matter. C-SEO Bench evaluated methods across tasks, domains, and models and found only three statistically significant positive ranking improvements among 54 tested cases; many settings showed negative effects. SAGEO Arena found that body-text optimization could degrade retrieval in its full-pipeline benchmark and that effects differed by stage.

Those studies do not predict the uplift for your pages. They show that treatment effects can vary by method, task, domain, model, and pipeline stage. A credible local experiment must allow zero and negative results.

A decision-ready readout

Report the primary estimate with an uncertainty interval, then show:

  • treated and control baselines;
  • absolute changes in both groups;
  • the relative estimate and its unit;
  • sample sizes at page, prompt, and observation levels;
  • missingness and exclusions;
  • pre-trend chart;
  • treatment compliance;
  • spillovers and concurrent changes;
  • secondary outcomes labeled as secondary;
  • robustness and sensitivity analyses;
  • the decision the result supports.

Avoid “the test worked” when the interval includes material harm and benefit. “The estimate is inconclusive under this sample and window” is a valid result. It tells the team what information is still missing.

Xindar’s public research specification requires frozen prompts, declared markets and languages, repeated observations, source IDs, separate dimensions, and versioned limitations before results are interpreted. A causal study adds assignment, controls, pre-periods, and a counterfactual model to that record.

Frequently asked questions

How long should a GEO experiment run?

There is no universal duration. Use baseline variability, expected effect, observation frequency, and the number of independent units to plan the study. Do not stop because a favorable week appears.

Can we use competitors as controls?

Sometimes, but competitors receive unknown treatments and may face different demand. Demonstrate comparable pre-trends and record major competitor changes where observable.

Is a before-and-after comparison ever enough?

It can describe change. It supports a causal claim only under strong assumptions that unrelated time-varying causes did not affect the outcome.

Should traffic or leads be the primary outcome?

Only when volume and attribution support it. Citation or support outcomes are closer to the intervention; leads are farther downstream and usually require larger samples.

Source and method note

This article draws on Callaway and Sant’Anna’s difference-in-differences framework, Brodersen et al.’s Bayesian structural time-series method, C-SEO Bench, SAGEO Arena, and Xindar’s public protocol. Sources were reviewed on September 9, 2026. The 40-page example and all percentages in the worked illustration are hypothetical. No client experiment, platform result, or uplift claim is reported.

Back to insightsMarkdown version