All insights
XINDAR INSIGHT

The Anatomy of a Cited Question: Why Only Some Queries Ever Reach the Retrieval Layer

Full-sentence questions trigger AI answers; two-word queries rarely do. The anatomy of which questions get retrieval, which get memory, and which get cited.

Every AI search strategy conversation eventually reaches the same unexamined assumption: that all queries are equally likely to produce an AI answer with citations. They are not. The gap between query types is not a nuance — it is the difference between a question that activates live retrieval (where citations happen) and a question the model answers entirely from memory (where your recent work does not exist).

Understanding this anatomy matters because it changes what "optimization" means. You cannot be cited by an answer that never consulted the index. And per the research below, a large share of queries never do.

Finding One: Query Structure Predicts Trigger Rates

Pew Research Center's behavioral study — 900 US adults with browsing trackers, 68,879 unique Google queries — captured precisely which searches produced AI Overviews, and the structural pattern is stark:

Query characteristicAI Overview trigger rate
Who/what/when/why question phrasing60%
Ten or more words53%
Full-sentence phrasing36%
All queries (average)~18%
One- to two-word queries8%

Read the spread: a question-shaped, ten-word query is roughly seven times more likely to summon an AI answer than a two-word keyword. The queries that trigger the answer layer are the ones that look like questions — because they are expensive to answer with a link list, and cheap to answer with synthesis. Short navigational and transactional queries mostly still get the classic results page (and, per the vertical data covered earlier in this series, transactional categories get actively protected from AI answers).

Ahrefs' keyword-level research points the same direction from the other side: 99.2% of keywords that trigger AI Overviews are informational in intent. The answer layer is, overwhelmingly, an informational layer.

Finding Two: Retrieval Is a Trigger, Not a Default

The deeper anatomy question is what happens inside the AI system when a query arrives — and here the research distinguishes two operating modes that most strategy content conflates.

Model-memory mode: for evergreen informational questions — definitional, conceptual, stable knowledge — assistants frequently answer entirely from parametric memory. No index is consulted, no sources are retrieved, nothing you published last month can possibly appear. Moz's fan-out research formalizes the distinction as In-of-Model versus Out-of-Model responses — visibility that comes from being in training data versus visibility that comes from live retrieval. These are two different games with two different clocks: training-data presence accrues over years; retrieval presence can move in weeks.

Retrieval mode: certain query characteristics reliably switch systems into search-and-cite behavior. ChatGPT's citation behavior, as documented in citation studies, triggers retrieval for timely queries, factual queries, and comparison queries — pricing, news, "best X for Y" — while evergreen informational questions often get answered with zero citations. The commercially loaded queries are precisely the ones that reach for the index.

The strategic implication deserves its own sentence: the queries most likely to retrieve and cite are the queries closest to a purchase decision. That is good news for brands — the answer layer concentrates its引用 behavior exactly where commercial intent lives — and a hard filter for what content work actually matters.

Finding Three: Fan-Outs Are Where Brands Live

Google's AI Mode decomposes a single user question into multiple internal sub-queries — the "query fan-out" — and answers from the synthesis. Moz's 2026 dataset (50,000 fan-out prompts across 20 industries, plus 5,333 Gemini grounding queries) mapped this hidden query space, and the geography matters:

  • Entity fan-outs ("what is [brand]") and comparison fan-outs ("[brand] vs [alternative]") account for 97% of all brand mentions in AI answers

  • Fan-out prompts average 8.2 words — a middle-tail of natural-language variants, far too many to track individually

  • The alignment between fan-out prompts and their topic clusters averages 0.67 cosine similarity — related, but not redundant

Two consequences follow. First, the visible query a user types is a proxy for an invisible cluster of internal questions; optimizing for the typed keyword is optimizing for the tip of the fan-out iceberg. Second, brand visibility concentrates in exactly two fan-out types — which means coverage of entity and comparison query space is not one task among many; it is the task.

This is also where the "keyword research is dead" debate lands. Sistrix's prompt research — built on 62 million real user questions — argues that keyword-level research no longer describes AI search behavior, and the fan-out evidence supports the softer version of that claim: the unit of optimization is shifting from the keyword to the question and its fan-out cluster. Keyword tools still matter for volume; question-level research (what buyers actually ask, in their words, per platform) is where citation opportunity is now visible.

Finding Four: Grounding Changes the Answer Mid-Conversation

One more mechanism completes the anatomy. Gemini distinguishes grounding queries — where the model reaches for real-time facts mid-response — from answered-from-memory generations. Moz's dataset isolated 5,333 of them. The practical meaning: whether your content participates in an answer can vary within a single conversation, as the assistant decides question-by-question whether memory suffices or the index is needed.

For brands, this produces the two-clock problem from Finding Two in operational form:

  • The training-data clock (years): Wikipedia presence, reference-layer mentions, wide earned-media coverage — the System Memory layer from the GEO Stack framework

  • The retrieval clock (weeks): fresh, extractable, well-structured content that wins the next retrieval wave

A strategy that only feeds one clock leaves the other running unmanaged.

What This Means for Question Research

Synthesizing the anatomy into a working method:

  1. Filter your topic list by trigger probability. Question-shaped, sentence-form, ten-word-plus, informational-comparison queries are the answer layer's home turf. Two-word transactional queries mostly are not — spend accordingly.

  2. Build your prompt set around fan-out types, not keywords. Entity and comparison clusters carry 97% of brand mentions; a tracking set without both is measuring the wrong universe (per the measurement discipline from earlier in this series).

  3. Write for retrieval mode. Facts with dates, pricing and specification clarity, quotable comparisons — the content characteristics that make a page usable when a "best X for Y" query triggers live search.

  4. Accept the memory game as a separate investment. Wikipedia adjacency, reference coverage, and community ubiquity operate on the training-data clock and cannot be sprinted. Schedule them as infrastructure.

  5. Test in the user's phrasing. Because trigger behavior is phrasing-sensitive (36% for sentences vs 8% for short queries), prompt-set language matters as much as prompt-set topics.

The Honest Caveats

  • Pew's trigger data reflects Google's 2025 AI Overview behavior — trigger policies are product decisions that shift (per the vertical patchwork piece, rates moved dramatically within a year). The structural gradient (questions trigger, short queries don't) is likelier to persist than the specific percentages.

  • Trigger behavior is not citation behavior. A triggered answer may still cite nothing, or cite competitors. This anatomy explains reach; the citation mechanics — source profiles, formats, freshness — are covered in earlier pieces.

  • Fan-out internals are inferred, not disclosed. Moz's dataset simulates fan-outs with LLMs; Google does not publish its actual decomposition. Treat the 97% concentration as the best available map of a hidden territory.

The Bottom Line

The answer layer is not uniformly laid over the query universe — it is a patchwork that activates for question-shaped, information-hungry, commercially-loaded queries and stays dormant for the rest. Within the activated territory, brand visibility concentrates overwhelmingly in entity and comparison question clusters, and participation depends on which of two clocks — retrieval or memory — the query invokes.

The teams that internalize this anatomy stop asking "how do we rank for keywords" and start asking a better question: which questions trigger the retrieval layer, what gets cited when they do, and are we present in the fan-out clusters our buyers actually generate? In the anatomy of a cited question, that is the whole game.


Sources

  • Pew Research Center (2025): 900 US adults, 68,879 unique queries; AI Overview trigger rates by query structure (question phrasing 60%, 10+ words 53%, sentences 36%, 1–2 words 8%); ~18% average trigger rate

  • Ahrefs AI Overviews research (2025): 99.2% of AIO-triggering keywords are informational intent (300,000-keyword study)

  • ChatGPT citation trigger analysis (2025–2026 citation studies): retrieval triggered by timely/factual/comparison queries; evergreen informational questions often answered with zero citations

  • Moz, "What 50k Query Fan-Outs Reveal About Brands" (2026): entity + comparison fan-outs = 97% of brand mentions; 8.2-word average; 0.67 alignment; In/Out-of-Model response distinction; 5,333 Gemini grounding queries

  • Sistrix Prompt Research (2026): 62 million real user questions; keyword-research obsolescence thesis

  • The GEO Lab GEO Stack framework (2026): System Memory layer (the training-data clock)

  • Cross-context from this series: The Measurement Trap (prompt-set design), A Patchwork Not a Layer (vertical trigger rates), The Vanishing Click (Pew methodology details)

Trigger percentages are snapshots of product behavior as of the study dates; engines re-tune continuously. Fan-out mechanics are inferred from simulated datasets, not official disclosures.

Back to insightsMarkdown version