All insights
XINDAR INSIGHT

Beyond the Citation: Did Your Page Actually Shape the Answer?

Learn how to separate citation selection, claim support, observable source contribution, and causal dependence.

By:Daoyu Guan, Head of GEO Operations & Editorial Lead at Xindar
Reviewed: September 8, 2026

Direct answer: A citation proves that an answer displayed a source, but it does not prove that the source supplied the answer's central facts, supported the sentence beside the citation, or caused the answer to take its final form. A defensible GEO review separates source selection from source contribution, breaks the answer into atomic claims, checks support claim by claim, and labels textual influence as an observational proxy unless a counterfactual test shows that removing the source changes the answer.

The easiest GEO result to count is often the easiest to overstate.

A dashboard records that a page was cited. Marketing reports a visibility win. The source appears next to a paragraph, so the team assumes that its research shaped the answer. Several different events have been compressed into one metric.

The system may have selected the page as a background reference while using another source for the decisive number. It may have paraphrased the page closely but attached the citation to a sentence containing an extra, unsupported claim. It may have cited several pages and synthesized a statement that none of them supports on its own. The citation is real in every case. Its evidential role is different.

Two research lines help clarify the problem. The 2023 paper “Evaluating Verifiability in Generative Search Engines” distinguishes citation completeness from citation correctness. In a human evaluation of four systems available at the time, only 51.5% of generated sentences were fully supported by citations on average, and 74.5% of citations supported the associated sentence. Those figures are a historical audit of named 2023 products, not current error rates.

A 2026 pre-submission study, “From Citation Selection to Citation Absorption”, proposes another distinction: whether a platform selects a source, and how deeply the source appears to contribute language, evidence, structure, or factual support to the answer. Its public dataset contains 602 controlled prompts across ChatGPT, Google AI Overview/Gemini, and Perplexity, more than 21,000 valid search-layer citations, and 18,151 successfully fetched pages. The reported patterns are descriptive and the manuscript explicitly avoids claiming access to hidden model internals.

Together, these papers suggest a better reporting question. Do not ask only, “Were we cited?” Ask, “What role did this source play, and how well does it support the claim?”

Four observations hidden inside one citation

ObservationQuestionUnit to inspect
SelectionDid the interface list this URL as a source?answer-source pair
SupportDoes the source justify the attached claim?atomic claim-source pair
AbsorptionDoes the answer appear to use this source's facts, language, or structure?answer-source relationship
DependenceWould the answer materially change without this source?counterfactual answer pair

Selection is visible. Support requires reading. Absorption is inferred from observable correspondence. Dependence requires an intervention or a credible natural experiment.

These levels should not be reported as synonyms. A source can be selected without supporting the adjacent sentence. A source can support a claim without being the only plausible origin of it. Strong textual overlap can suggest use, but common wording may occur in many documents. Even a counterfactual change must be interpreted carefully because generative systems vary across runs.

Citation correctness and completeness answer different questions

Suppose an answer says:

Acme's Model R pump handles 180°C continuously, is certified to ISO 9001, and ships to Germany in five days.[1]

This is one sentence but at least three claims:

  1. Model R has a continuous operating limit of 180°C.
  2. A relevant organization or site holds an ISO 9001 certificate.
  3. The product ships to Germany within five days.

Source [1] may support the temperature under a named test condition, show a certificate for a different legal entity, and say nothing about current lead time. Marking the whole sentence “cited” hides partial support.

Citation correctness asks whether the cited source supports the claim attached to it. Citation completeness asks whether claims that need support have support. A response can score well on one and poorly on the other. It may attach accurate sources to half its factual claims, or attach a source to every sentence even when several links are irrelevant.

For high-stakes industrial, health, financial, or legal facts, sentence-level review may still be too coarse. Break a compound sentence into atomic claims that can be judged separately.

What “citation absorption” tries to measure

The 2026 absorption paper treats source contribution as a second outcome after selection. Its public influence score combines repeated reference, early appearance, coverage across answer paragraphs, TF-IDF similarity, and bigram or trigram overlap. This is useful because it moves analysis from a source list toward the relationship between page and answer.

It is still a constructed proxy. The score does not reveal model attention, a private retrieval rank, or causal dependence. Its components also cannot be turned around and presented as independent proof of what caused a high score; they partly define the score itself.

The paper reports a striking descriptive contrast. In its dataset, Perplexity averaged 16.35 citations per prompt, Google 12.06, and ChatGPT 6.88. Among successfully fetched pages, ChatGPT had the highest mean influence score: 0.2713, compared with 0.0584 for Google and 0.0646 for Perplexity. Fewer displayed sources coexisted with deeper average textual influence in this sample.

That is why citation count alone is ambiguous. A broad answer may distribute attention across many references. A narrower answer may lean heavily on a few. Neither pattern is automatically better. The right metric depends on whether the publisher wants exposure, referral opportunity, factual attribution, or substantive contribution.

A practical claim-support coding system

Use five labels. They are simple enough for an editorial team and precise enough to audit.

LabelMeaningExample decision
Fully supportedSource justifies the whole atomic claim, including scopeexact model, condition, unit, and date match
Partly supportedSource supports only part or a weaker versionmaximum value shown, continuous-use condition absent
ContradictedSource provides incompatible informationcertificate expired or covers another facility
UnrelatedSource does not address the claimcorporate history cited for delivery time
Not assessableSource is unavailable or evidence is ambiguousblocked page, missing attachment, unclear table

The reviewer should preserve the answer, the cited URL, the relevant source passage, the decision, and a short reason. “Looks right” is not a reason. “Source names Model R but gives 180°C only as a short-duration maximum” is.

How to test whether a page shaped an answer

  1. Freeze the observation. Save the exact prompt, platform, date, market, language, account context, answer, and cited URLs. A later run is a new observation.

  2. Segment the answer into atomic claims. Separate entities, numbers, dates, relationships, recommendations, and conditions. Keep evaluative language distinct from factual statements.

  3. Map citations to claims. Interfaces often place one citation after a long paragraph. Record every claim the placement appears to cover, then judge each source against each claim.

  4. Locate the supporting passage. Copy only the necessary source fragment and its context. Record whether the source is first-party, official, academic, regulatory, journalistic, or another class.

  5. Code support quality. Apply the five labels consistently. Use a second reviewer for consequential claims and record disagreements instead of forcing consensus silently.

  6. Estimate observable absorption. Look for distinctive facts, uncommon phrasing, sequence, examples, and table structure shared by source and answer. Label this as correspondence, not hidden causation.

  7. Run a counterfactual where possible. In a controlled retrieval system, generate with and without the source while holding the prompt and other context fixed. In a public system, repeated runs or temporary availability changes are weaker because the platform can change other inputs.

  8. Report the right outcome. Separate selection rate, support rate, observed contribution, and answer accuracy. Attach sample size and dates.

Why text similarity is not enough

Text overlap is attractive because it is cheap. It can also mislead.

A definition may be standardized across dozens of sources. A product specification may be copied by distributors. A press release may be syndicated verbatim. High overlap shows correspondence with the page you measured; it does not establish that the engine used that page rather than another copy.

Low overlap is not proof of low contribution either. A system can compress a table, translate a Chinese passage into English, combine two paragraphs, or preserve the facts while changing every sentence. Semantic similarity can detect some of this, but it remains an inference.

The strongest observational case combines several signals: the cited URL is present, the answer contains a distinctive and correctly scoped fact, the fact appears on the page, alternative sources do not obviously contain it, and the relationship repeats across controlled observations. Even then, use “consistent with contribution” unless an intervention supports a causal claim.

Q&A formatting is not a contribution strategy

The absorption paper reports that pages classified as Q&A averaged an influence score of 0.0947, compared with 0.1005 for non-Q&A pages in its feature table. This descriptive result does not prove Q&A causes lower influence. It does challenge the rule that converting everything into questions and answers will automatically deepen source use.

The useful unit is an evidence-bearing answer, not the presence of a question mark. A Q&A entry that says “Yes, we offer high quality and fast delivery” contains little to absorb. A short technical section that names the product, states the condition, gives the method, and links the controlled record may contribute far more.

Use FAQ structure for real recurring questions. Do not use it as decorative markup around thin claims.

A reporting dashboard that does not overclaim

MetricDefinitionReview note
Source selection rateobservations where the URL is displayed as a source / eligible observationsreport prompt panel and platform
Claim support ratefully supported atomic claims / cited atomic claims revieweddo not combine “partial” silently
Citation precisionsupporting citations / citations revieweddefine the attachment rule
Citation completenesssupport-needing claims with adequate citation / such claims reviewedrequires a claim policy
Observable contributionshare of answers with distinctive factual or structural correspondenceproxy, not model attention
Answer accuracycorrect atomic claims / claims reviewedcan improve while citations fall, or vice versa
Qualified referralvisits or inquiries attributable under stated collection rulescommercial outcome, not epistemic quality

Xindar's public measurement protocol similarly separates mentions, citations, recommendations, accuracy, source quality, and commercial signals. The important feature is not the vendor name. It is the refusal to collapse unlike observations into one score.

Frequently asked questions

If our page is cited, can we say the AI trusts us?

No. You can say the page was cited in the recorded answer. “Trust” is a broader interpretation that would require a definition and stronger evidence.

Can a source shape an answer without being cited?

It is possible in principle, but a public observer usually cannot establish it. Similar wording may come from model memory or another source. Report the uncertainty.

Should we optimize for more citations or deeper use?

Choose according to the decision. Publishers seeking referral opportunities may value selection breadth. Technical brands may care more about correct use of approved facts. A mature program measures both.

Are the 2023 support rates still current?

They should not be presented as current platform error rates. They are a historical human audit that demonstrates why citation presence and support must be evaluated separately.

Source and method note

This article relies on Liu, Zhang, and Liang's verifiability study, the 2026 citation absorption pre-submission draft, and Xindar's public research method. The 2026 manuscript labels its numerical results as descriptive and specifies additional work needed for confirmatory inference. Xindar's private knowledge-base records informed the governance workflow but are not treated as independent evidence. No live platform experiment was conducted for this article.

Back to insightsMarkdown version