All insights
XINDAR INSIGHT

Live, Cached, or Closed: The Retrieval Setting Behind an AI Answer

An AI answer can look like a judgment about the web when it is partly a consequence of how the application was configured.

An AI answer can look like a judgment about the web when it is partly a consequence of how the application was configured. A live search tool can retrieve a page changed this morning. A cached search tool may work from an earlier copy. A domain allowlist can exclude almost every source before ranking begins, while a disabled search tool leaves the model to answer without current web evidence. GEO measurement therefore needs to record the retrieval environment, not just the prompt and final citation. Otherwise, teams may attribute a visibility change to content quality when the real cause is freshness, location, source policy, or the absence of search.

The September update that made the hidden variable visible

OpenAI introduced its Agents API in public beta on September 10, 2026. The accompanying web-search documentation exposes several controls that are easy to overlook when evaluating AI visibility: search can run in live, cached, or disabled mode; developers can restrict results to as many as 100 allowed domains; they can supply a location; and they can change the amount of search context passed to the agent. The documentation also states that merely asking an agent to search does not enable search when the tool has not been configured.

This is product documentation for one API, not a universal description of every AI answer engine. Its broader importance is conceptual. It shows that a conversational answer is produced inside a retrieval environment. Two applications can use the same model and receive the same question while operating over different evidence. Even the same application can produce different source sets after a configuration change.

The useful unit of GEO analysis is therefore not "the model." It is the complete answer system at a particular time.

A more accurate model of an answer

A conventional visibility report often records four fields: platform, prompt, answer, and cited URLs. That is not enough for a causal diagnosis. A more complete representation is:

observed answer = f(model, conversation, tool availability, retrieval mode, corpus time, domain policy, location, context budget, answer policy)

Each term can change the result.

VariableWhat it controlsTypical false conclusion when omitted
Tool availabilityWhether current web evidence can be fetched"The model does not know our new page"
Retrieval modeLive internet, saved content, or no search"Our update was ignored"
Corpus timeThe latest version available to retrieval"The old claim still ranks better"
Allowed domainsWhich sources may enter the candidate set"Competitor A is universally preferred"
LocationGeographic interpretation and local source selection"Results are globally consistent"
Context sizeHow much retrieved material reaches generation"Every fetched source influenced the answer"
Conversation stateConstraints inherited from earlier turns"The prompt is reproducible"

This table does not imply that every consumer product exposes every variable. It establishes the variables an evaluator should control or mark as unknown.

Live search is not the same as immediate visibility

The word "live" can be misleading. It means the tool may access the current internet; it does not promise that every URL is already discovered, fetchable, parsed, selected, or cited. A newly edited page must still pass through several gates:

  1. The system must know the URL or discover it from a query.
  2. The page must be accessible to the relevant fetch path.
  3. Its useful content must be present in a representation the system can process.
  4. The page must be judged relevant enough to enter the working context.
  5. Its claim must survive synthesis alongside other evidence.
  6. The interface must decide whether and how to expose a citation.

A failed appearance at any one gate can produce the same visible result: no mention. That is why publishing at 10:00 and checking an AI answer at 10:05 is not a valid indexation test. It combines discovery latency, fetch behavior, retrieval selection, and generation into one observation.

Cached search creates a version problem

Cached retrieval is useful for speed, repeatability, cost control, and resilience. It also changes what "current" means. Consider a software vendor that changes its retention policy from 90 days to 30 days. The website now shows 30 days, but a saved copy still contains the old number. A cached agent can accurately quote the evidence available to it and still give an outdated answer.

The appropriate remedy is not to repeat the new number across dozens of pages. That can create internal contradictions and make future maintenance harder. Instead, maintain one canonical policy page, add an explicit effective date, update dependent pages, preserve a concise change record, and test whether the old value remains in feeds, PDFs, help-center mirrors, partner listings, or translated versions.

For volatile facts, a useful content block states five things together:

  • the value;
  • the entity or plan to which it applies;
  • the effective date;
  • the jurisdiction or audience;
  • the canonical source.

That compact bundle is easier to retrieve and safer to quote than a promotional sentence detached from scope.

Domain allowlists change the meaning of "best source"

An allowed-domain list does more than favor certain publishers. It defines the evidence universe before source selection. If an enterprise research agent searches only government and standards domains, a vendor's excellent explainer may be ineligible. If a customer-support agent searches only the company's documentation, independent reviews will never enter the answer. If a shopping assistant uses approved retailers, an editorial comparison site may disappear from consideration.

This does not make the answer defective. Domain restrictions can be an appropriate safety and governance control. It does mean that a citation win inside one application cannot be generalized to all AI search.

GEO teams should classify tests into at least three source regimes:

RegimeExampleWhat the result can support
Open webNo known domain restrictionBroad discovery under the tested platform conditions
Curated webA documented allowlist or approved-source setVisibility inside that governed corpus
Owned corpusCompany documentation or private knowledge baseFindability and answer quality within first-party material

Reporting all three as "AI ranking" erases the most important difference between them.

Location is part of the query even when the user never types it

The Agents API can receive country, region, city, and time zone as search-location hints. Other systems may infer geography from account settings, device state, interface language, or network information. A prompt such as "Which provider is available near me?" is incomplete without that context, but even apparently global questions can have local answers because product availability, regulation, language, and source access vary by market.

A defensible visibility study records both the explicit prompt and the environmental location. For international brands, it should use a market matrix rather than one global average. A source that performs well in London but is absent in Singapore is not unstable in the statistical sense; it may be responding to different source and product conditions.

A six-run diagnostic for retrieval-sensitive questions

When an answer appears outdated or omits an important source, use a controlled sequence.

  1. Run the baseline. Save the exact prompt, prior conversation, timestamp, account state, location, answer, links, and screenshots or raw output available through the product.
  2. Ask a freshness probe. Use a fact that changed recently and has a clear official date. This does not prove live retrieval, but it can reveal an obviously stale path.
  3. Name the canonical source. Ask the system to consult the official page. Compare whether the claim changes and whether the page is cited.
  4. Constrain the domain where the product allows it. In an API test, use an official-domain allowlist. This separates source discovery from source interpretation.
  5. Change one location variable. Keep the prompt and all other settings fixed. Record whether the source set and conclusion change.
  6. Repeat over time. Run the same panel after the suspected cache or discovery interval. A single successful retry is evidence of availability, not stable visibility.

The sequence produces a diagnosis such as "the page is retrievable when named, but not selected in open-web search" or "the live and cached runs disagree on the policy date." Those statements are actionable. "The AI does not like our site" is not.

Design content for both current and saved retrieval

Teams cannot control an application's cache policy, but they can reduce version ambiguity.

First, keep stable canonical URLs for durable facts. Replacing a policy page with a new path at every update fragments historical references. Second, place the last-reviewed or effective date near the claim it qualifies. Third, expose the decisive text in the page itself rather than only in an image, animation, or downloadable attachment. Fourth, maintain clear entity and product identifiers so an older variant is not confused with the current one. Fifth, make corrections explicit. A short note such as "Updated September 18, 2026: the Pro plan now includes 30-day retention" gives a retriever evidence that the value changed.

Do not manipulate visible dates to simulate freshness. A changed date without a changed fact weakens auditability and can make readers distrust legitimate updates.

Measure the pipeline, not one output

A retrieval-aware GEO dashboard can use five stages:

StageQuestionEvidence
DiscoveryCan the system find the URL for the target query?URL or domain appears in retrieved/cited set
FreshnessWhich version of the fact is used?Value plus effective date
SelectionDoes the source enter the answer context?Citation, link, tool trace, or controlled source test
SynthesisIs the claim represented accurately and with its conditions?Claim-level review
StabilityDoes the result persist across runs and locations?Repeated panel with fixed configuration

Where a consumer interface hides tool traces, mark selection as unobserved rather than inferred. A citation proves that a source was exposed to the user. It does not reveal every source retrieved or every rule used to select it.

What this changes for GEO strategy

The immediate lesson is methodological. Content work, technical access, and measurement must share the same retrieval model. An editorial team can publish an excellent update while an evaluation team tests a cached corpus and reports failure. A developer can add a domain allowlist for safety while marketing interprets the resulting source concentration as a brand preference. A regional team can improve local pages while a global dashboard averages away the gain.

Before changing content, identify which part of the pipeline is failing. If a source cannot enter the candidate set, rewriting the introduction will not fix the problem. If the page is retrieved but the claim lacks scope or date, clearer evidence may help. If the answer is correct but uncited, the issue is attribution or interface exposure, not factual visibility.

Frequently asked questions

Does live search guarantee that a newly published page can appear?

  1. Live access permits current web retrieval, but discovery, access, parsing, relevance, synthesis, and citation are separate gates.

Is a cached answer necessarily wrong?

  1. Cached retrieval can be appropriate for stable facts and reproducible workflows. It becomes risky when the question depends on recent prices, policies, availability, or events.

Can a prompt force an agent to search the web?

Not always. OpenAI's Agents API documentation says that web search must be configured as a tool; asking for search in the prompt does not enable a missing tool.

Should GEO reports compare models or applications?

Compare complete applications under recorded conditions. Model name alone does not identify retrieval mode, source restrictions, location, or interface policy.

What is the fastest test for a suspected stale answer?

Ask for the official source and effective date, then compare a live controlled-domain run with the original environment where possible. Treat the result as a diagnosis of that setup, not a universal platform finding.

Sources and evidence boundary

The product settings described here are documented OpenAI API behavior as of the access date. The pipeline model and audit procedure are analytical recommendations. They do not claim that all AI products use the same retrieval architecture or that any configuration guarantees inclusion, ranking, citation, or traffic.

Back to insightsMarkdown version