All insights
XINDAR INSIGHT

The Citation That Proves Nothing: How to Audit Whether a Source Supports an AI Claim

A citation supports an AI claim only when the cited source entails the claim at the same scope: the same entity, relationship, value, unit, condition, place, and time. A working link proves access to a page.

Direct answer: A citation supports an AI claim only when the cited source entails the claim at the same scope: the same entity, relationship, value, unit, condition, place, and time. A working link proves access to a page. It does not prove that the page supports the adjacent sentence. Audit the claim-source pair, record full, partial, contradictory, absent, or unassessable support, and evaluate coverage separately from citation accuracy.

The most misleading citation is often a respectable one.

A government page, peer-reviewed paper, or manufacturer specification can be authoritative and still fail to support the sentence beside it. The source may discuss the wrong jurisdiction, measure a different outcome, describe an older product revision, or support only half of a compound claim. Authority answers “How credible is this source for its domain?” Support answers “Does this source justify this exact claim?”

GEO teams need both questions. Counting links measures source display. Auditing support measures whether the answer is verifiable.

Citation presence, coverage, and correctness are different

Three observations are commonly collapsed into “citation quality”:

ObservationQuestionWhat a positive result establishes
Citation presenceIs an inspectable source displayed?A source link was attached or listed
Support coverageAre claims that require evidence supported by citations?Fewer verification-worthy claims are left unsupported
Citation correctnessDoes each cited source support the claim it is attached to?Fewer irrelevant or misleading source attachments

The 2023 study Evaluating Verifiability in Generative Search Engines formalized related measures as citation recall and citation precision. Recall asks what proportion of verification-worthy statements are fully supported by their associated citations. Precision asks what proportion of citations support their associated statements.

Those names can confuse marketing teams because “recall” and “precision” are also used in retrieval. Preserve the paper’s definitions in research reports, but include plain-language labels such as support coverage and citation correctness.

Why a URL is not evidence of entailment

Entailment is the core relationship. If the source is accepted as true, would a careful reader agree that the claim follows?

Suppose an answer says:

Supplier K’s AX-4 is CE certified, operates continuously at 180°C, and can be delivered worldwide in 21 days. [1]

The linked page may say that the manufacturer applies CE marking to one regulated product line, reports a 180°C short-duration test for another model, and gives a domestic planning estimate of 21 days for standard orders. The page contains all the important words. It does not support the answer.

The claim fails on at least four dimensions: product identity, meaning of the conformity statement, duration of the temperature condition, and geographic scope of the delivery estimate. A keyword-overlap check could pass it. A support audit should not.

Compound sentences hide partial support

Sentence-level review is often too coarse. One sentence may contain several atomic claims joined by “and,” a comma, a relative clause, or an implied relationship.

Break the example into these units:

  1. Supplier K offers product AX-4.
  2. AX-4 is within the relevant CE-marked product scope.
  3. AX-4 operates continuously at 180°C.
  4. Supplier K delivers AX-4 worldwide.
  5. The delivery time is 21 days.

Each unit needs a judgment. The ALCE benchmark paper similarly evaluates whether statements are fully supported and whether individual citations are relevant. Its human evaluation asked annotators whether a citation fully supports, partially supports, or does not support a sentence. The paper also notes a limitation of automated natural-language-inference scoring: partial support can be difficult to detect.

For business review, five labels are more useful:

LabelDecision ruleReporting treatment
Full supportThe source justifies the entire atomic claim at the stated scopeCount as supported
Partial supportThe source justifies a weaker, narrower, or incomplete versionReport separately; do not count as full
ContradictionThe source states something incompatible with the claimEscalate and correct
No supportThe source is relevant to the topic but does not justify the claimCount as unsupported
UnassessableThe page is unavailable, ambiguous, or requires inaccessible contextPreserve uncertainty; do not assume support

“Partial” must include a reason. Useful reason codes include wrong date, narrower market, different model, missing condition, missing denominator, unsupported causal language, and source only supports an input rather than the conclusion.

Multiple citations can support one claim together

One source does not always need to prove everything. A claim may be supported by a union of citations.

For example, an official registry may establish a company’s legal identity, a certification-body directory may establish certificate status, and a product declaration may connect the certified entity to the product. Together they can support a carefully worded statement. None may be sufficient alone.

Union support should satisfy three rules:

  • every step in the reasoning chain is explicit;
  • the sources refer to compatible entities, dates, and scopes;
  • no source contradicts another material part of the claim.

Do not use a pile of loosely related links as a substitute for a missing premise. Ten sources that discuss CE marking in general do not establish that one named product meets the relevant requirements.

Historical research shows why manual review matters

The 2023 verifiability study audited Bing Chat, NeevaAI, Perplexity, and YouChat across several query sets. Averaged across that study, 51.5% of generated sentences were fully supported by citations, and 74.5% of citations supported their associated sentence. The authors also found that results varied by system and query distribution.

These figures describe a historical study. They are not 2026 error rates for current products, and they should not be used to rank today’s platforms. Their continuing value is methodological: fluent answers can look useful while support coverage and citation correctness remain incomplete.

The study also found a negative correlation between perceived utility and citation precision in its evaluated systems. The authors hypothesized that close copying could raise support while producing less useful answers. That result warns against another shortcut: perfect support does not by itself make an answer relevant, complete, or well reasoned.

A claim-by-claim audit protocol

1. Freeze the answer

Save the exact prompt, answer, citations, platform or product, date, market, language, account context, and visible model label. A rerun is a new observation because sources and wording can change.

2. Define the attachment rule

Record which claim each citation appears to support. A superscript at the end of a paragraph may attach to the last sentence, the whole paragraph, or a source panel with no explicit mapping. Do not silently choose the interpretation that produces the best score.

3. Mark verification-worthy content

Names, dates, quantities, comparisons, causal claims, legal or regulatory statements, specifications, quotations, and claims about external events generally require verification. Pure transitions and clearly labeled opinions may not.

4. Decompose into atomic claims

Split compound sentences until each row can receive one support judgment. Preserve the original wording and assign stable claim IDs.

5. Capture the source passage

Save the shortest passage that could support the claim, plus enough context to interpret it. Record title, publisher, URL, publication or update date, and access date. For a PDF or table, preserve page, section, row, and unit.

6. Compare the full scope

Check entity, attribute, polarity, quantity, unit, denominator, method, condition, geography, date, and uncertainty. Also check the direction of reasoning. Evidence that A is associated with B does not support “A caused B.”

7. Assign support and source-quality labels separately

A low-quality source may fully support what it says while remaining a poor authority. An official source may have no support for the attached claim. Keep support_status and source_class in separate fields.

8. Use a second reviewer for consequential claims

Regulatory, safety, medical, financial, contractual, and high-value product claims deserve independent review. Record disagreements and the adjudication reason.

The minimum audit record

FieldExample
Observation ID20260909-Q07-BING-R2
Claim IDC04
Exact claim“AX-4 operates continuously at 180°C.”
Citation attachmentsource [2], paragraph-end placement
Evidence passage180°C short-duration maximum, 30 minutes
Support statusContradiction
Reason codeduration mismatch
Source classfirst-party test report
Reviewerinitials and date
Corrected claim“AX-4 Rev. C was tested at 180°C for up to 30 minutes under M-17.”

This record is auditable because another reviewer can reconstruct the decision. A dashboard percentage without these rows cannot reveal whether disagreements came from source quality, attachment ambiguity, compound claims, or reviewer inconsistency.

Calculate metrics without hiding the denominator

Use at least two rates:

Support coverage = fully supported verification-worthy claims ÷ verification-worthy claims reviewed.

Citation correctness = citations providing full support ÷ citations reviewed.

Report partial support separately. If the research question permits union support, specify how it is counted. Always publish the number of answers, claims, and citations reviewed. A 90% rate based on ten claims is not equivalent to 90% based on ten thousand.

Do not mix query classes casually. A factual lookup, open-ended essay, product comparison, and legal question impose different evidence burdens. The 2023 study found that query distribution affected citation recall. A program that changes its prompt mix between months changes its measuring instrument.

How this differs from source contribution

Support asks whether the source justifies the claim. Contribution asks whether the source shaped the answer. These can diverge.

A source may support a standard definition that appears in hundreds of places; support is high, but unique contribution is uncertain. A page may strongly shape an answer through a distinctive example while the attached citation fails to support one added number. That is why source-absorption analysis and citation-support auditing should remain separate reports.

Xindar’s measurement specification uses separate fields for citation presence and citation support and does not approve a composite score without weighting and missing-value rules. The useful principle is general: do not let a single “visibility” score erase an evidence failure.

Frequently asked questions

Does an authoritative domain make a citation correct?

  1. Authority concerns source quality. Correctness concerns whether the source supports the attached claim at the stated scope.

Should partial support count as half?

Only if that scoring rule was defined before review and serves the decision. For public reporting, show partial support as its own category.

Can an automated entailment model replace reviewers?

It can assist triage, but ALCE documents limits around partial support. Use human review for ambiguous and consequential claims and audit the model’s errors on your domain.

Does a high citation-correctness rate prove the answer is accurate?

  1. Important claims may be uncited, sources may share the same error, and the answer may omit decisive qualifications.
Back to insightsMarkdown version