All insights
XINDAR INSIGHT

Volatility Is Not a Confidence Interval

When an AI answer changes across repeated runs, the change is observed volatility. A confidence interval is an estimate of uncertainty under a defined sampling and statistical model.

Direct answer

When an AI answer changes across repeated runs, the change is observed volatility. A confidence interval is an estimate of uncertainty under a defined sampling and statistical model. They are related but not interchangeable. Citation turnover can result from model sampling, retrieval changes, query variation, source availability, or platform updates; it does not automatically reveal the probability that a true visibility rate lies within a range. Report the prompt, repeat rule, time window, unit, and denominator, then separate observed turnover from sampling uncertainty. Do not attach a statistical interval to a small panel unless the design supports it.

The number that looked more precise than the study

A team runs the same 20 prompts three times. Its brand appears in 11, 8, and 13 responses. The dashboard reports “53.3 percent visibility plus or minus 6 percent.” The number has a familiar statistical shape, but what does the interval mean?

If the three runs are treated as 60 independent observations, the calculation may be misleading because the same prompts were repeated and the responses may share platform state. If the team intends to estimate a rate for a larger population of buyer questions, the 20 prompts may not be a representative sample. If the platform changed between runs, the variation is partly time drift rather than random sampling.

The observed range from 40 to 65 percent is a fact about these runs. A confidence interval requires a target population, a sampling design, and assumptions about dependence and error. The two should not be merged because both contain numbers.

Define the uncertainty before calculating it

Ask what is being estimated. It might be the proportion of a defined prompt population in which a brand is cited under one platform configuration. It might be the proportion of responses that contain a verified source link. It might be the rate of an answer preserving a required product condition.

Each target has a different denominator. “One response” can mean one prompt-run pair, one user session, or one platform answer. “Citation” can mean a link, a unique cited page, a domain mention, or a source that actually supports the claim. Write the unit before collecting the counts.

Next identify the source of variation. Repeated sampling from a defined prompt population creates one kind of uncertainty. Model stochasticity creates another. Platform updates and content changes create time variation. A prompt panel that changes every day adds query variation. These sources can interact.

The critical GEO survey describes GEO as a stochastic, partially observable pipeline and finds that the reviewed evidence does not establish a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior. That does not make measurement pointless. It means a result needs an explicit design and boundary.

Volatility is an observation, not a diagnosis

Citation turnover can be summarized as the proportion of repeated runs in which the cited set differs. If runs cite pages A, B, C and then A, C, D, one page turned over. The measure says the observed source set changed. It does not say whether D is better, whether B was irrelevant, or which hidden stage produced the change.

Answer wording can change while the citation set stays constant. Conversely, a source can remain linked while its role in the answer changes. Track source identity, claim support, and answer meaning separately.

The Bing AI Performance documentation defines citation and cited-page measures and says they do not indicate ranking, authority, or placement. Those definitions are useful for deciding what was observed. They do not turn citation turnover into a confidence interval.

Google's AI performance report also has specific aggregation rules, including property versus page grouping, canonical assignment, preliminary data, and reporting limitations. A time series from that report has its own measurement boundaries. It should not be pooled casually with a prompt-panel rate.

A small panel can answer a small question

A prompt panel is valuable when it represents a defined set of decisions. It can tell you how a page performed across those prompts on specified dates and platforms. It cannot automatically tell you how the brand performs across all possible user questions.

The Xindar research methodology is a disclosed sampling protocol for visibility research. The page may be inaccessible in one collection attempt, so any reuse of its details should be rechecked directly before publication. The broader principle remains: prompt, market, language, repeat count, and observation rule belong in the record.

If a panel is deliberately balanced across question types, its percentage is a panel result. Do not call it a market share or demand-weighted visibility rate unless the sampling weights justify that interpretation. A panel can oversample high-consequence questions without being statistically representative, and that can still be the right business design.

Use repeated runs deliberately

  1. Define the target population and unit of analysis.
  2. Freeze prompt wording, language, market, page version, and platform configuration.
  3. Decide whether repeats test model stochasticity, time drift, or both.
  4. Capture the complete answer, source links, date, and relevant annotations.
  5. Label citation, support, recommendation, and condition preservation separately.
  6. Calculate the observed rate and turnover from the declared denominator.
  7. Use an interval only when the sample and dependence assumptions are defensible.
  8. Repeat after platform or content changes and mark the break in the series.

The OpenAI evaluation guidance recommends task-specific evaluation, calibrated feedback, and continuous evaluation. It does not prescribe a confidence-interval method for GEO panels. Choose a statistical method appropriate to the data or report descriptive uncertainty without pretending to estimate more than the design supports.

If repeats are identical because the system is deterministic under the tested settings, that does not prove the underlying behavior is stable. It only describes the tested configuration. If each response is generated independently, the runs still may share retrieval indexes, cached results, or content state.

Illustrate denominators with a fictional panel

Suppose 25 fixed prompts are run four times, producing 100 prompt-run observations. The brand appears in 47 responses. The observed prompt-run citation rate is 47 percent under that panel and date. If 12 prompts cite the brand in at least one of four runs, the prompt-level reach is 48 percent. Those are different measurements.

The 12 of 25 figure asks whether the prompt ever produced a citation. The 47 of 100 figure asks how often a prompt-run pair produced one. Neither shows how many buyers encountered the brand or how many cited pages supported the answer.

These numbers are invented. Their purpose is to show why a dashboard needs a denominator label. A team should store raw prompt IDs, run IDs, labels, and source URLs so someone can recompute the metric.

Do not average percentages from panels with different prompt counts or definitions without using the underlying counts. Do not compare a 20-prompt panel in English US with a 20-prompt panel pooled across five markets as if the samples represented the same population.

Confidence intervals need a sampling story

A confidence interval is meaningful relative to a procedure that repeatedly samples and calculates intervals. It is not a general-purpose bracket around any unstable number. If the prompt panel is a convenience sample, the interval can describe uncertainty under repeated sampling from that constructed panel, but it does not solve selection bias.

Repeated answers from one prompt are also not automatically independent Bernoulli trials. If a platform's retrieval state persists, or if the same prompt is repeated in a short period, the observations can be correlated. Treating them as independent can make the interval look narrower than warranted.

If statistical estimation is not the goal, report the observed range, run-level rates, prompt-level reach, and exact sample. That is often more honest and more useful for an editorial decision. If an interval is reported, name the method, assumptions, unit, and scope in the methodology note.

Separate platform change from content change

When a rate moves, log content releases, canonical changes, crawler access, prompt changes, product changes, and platform reporting updates. A before-and-after difference is a signal, not a causal estimate.

The C-SEO Bench tests content optimization under multiple actors and reports that many rewriting methods were ineffective or negative in its setting. The SAGEO Arena examines stage-level effects and notes that a change can affect retrieval, reranking, and generation differently. These findings support controlled experimentation; they do not give a shortcut for attributing a volatile public answer to one edit.

Use a control page or holdout question set where the business and platform conditions make that possible. Keep in mind that AI systems can share sources and market changes can affect both groups. State the design limitations rather than calling a simple rise a causal lift.

For recurring reporting, publish a compact volatility card beside the headline metric. It should show the number of prompts, runs per prompt, percentage of prompt-run observations with a citation, percentage of prompts cited at least once, unique pages observed, source-set turnover, collection dates, and any platform or content change during the window. This is more informative than a single confidence badge because it reveals whether the result is broad, concentrated, or unstable.

Keep the card comparable over time. If the panel grows from 25 prompts to 100, mark the break and avoid drawing a smooth trend line across it. If a new market is added, show the pooled result and the market-specific results separately. A larger sample can improve coverage while making the number less comparable to the old one.

Frequently asked questions

How many repeats should I run?

There is no universal count. Choose it based on the question, stochasticity you need to characterize, platform cost, and the precision required for the decision. More repeats do not repair a biased prompt panel.

Can a confidence interval show whether an AI citation will remain?

Not by itself. Persistence depends on platform, retrieval, content, query, and time conditions. An interval around a sampled rate is not a guarantee for a future answer.

Is volatility always a bad sign?

  1. It can reflect a changing information environment or broad source competition. The problem is interpreting volatility without knowing what changed or whether the variation matters to the user's decision.

What should an agency hand over?

The prompt panel, repeat rule, raw captures, labels, formulas, changes in platform or content state, and the exact meaning of every metric. This lets a client distinguish a measured observation from an unsupported trend claim.

Source and method note

Sources were retrieved on September 21, 2026. The fictional panel counts illustrate denominator choices and are not market data. Research papers and platform documentation are used within their stated scopes. No statistical estimate, customer visibility result, or named human review was produced for this article.

Back to insightsMarkdown version