All insights
XINDAR INSIGHT

A Citation Is Not a Recommendation: Labeling What an AI Answer Actually Says

An AI answer can cite a page without mentioning its brand, mention a brand without recommending it, recommend it only under stated conditions, warn against it, or exclude it from consideration. These are different outcomes and need different labels.

Direct answer

An AI answer can cite a page without mentioning its brand, mention a brand without recommending it, recommend it only under stated conditions, warn against it, or exclude it from consideration. These are different outcomes and need different labels. A reliable annotation system records the question, answer span, entity, statement type, stance, conditions, citation target, and whether the citation supports that exact statement. Annotation should occur at the claim or clause level, with uncertain cases preserved for review. Citation count measures source display. It does not by itself measure endorsement, rank, authority, factual support, or influence on a decision.

The dashboard says “cited”; the answer says “avoid”

Suppose an answer compares three payroll platforms. It recommends Platform A for multinational teams, calls Platform B suitable only for small domestic firms, and warns that Platform C lacks a required integration. A citation to Platform C's documentation supports the warning. A dashboard that counts all three citations as positive visibility erases the decision meaning of the answer.

The opposite problem also occurs. An answer may recommend a brand based on general knowledge or on sources displayed elsewhere, while the brand's own page is never cited. A citation-only report would miss the recommendation. Mention, source attribution, factual support, and recommendation are related events, but they are not substitutes.

This distinction matters for editorial strategy. A team responding to missing citations might improve source clarity. A team responding to repeated negative recommendations needs to inspect product fit, stale facts, or incorrect answer content. One blended “share of voice” number sends both teams to the same remedy.

Define the unit before defining the labels

The unit of annotation should be an entity-claim within a question-answer observation. The observation records the prompt, platform, market, language, date, run, and captured answer. Within that answer, annotators mark the smallest span that expresses a meaningful proposition about the entity.

A whole-answer label is too coarse. One response can recommend a product for one persona and reject it for another. A sentence can contain a positive feature and a material limitation. Clause-level annotation keeps those conditions attached.

The unit also needs a stable entity ID. “Acme,” “Acme Cloud,” and a product nickname may refer to one company, a product family, or different offerings. Resolve the entity before computing rates. Preserve ambiguous references as ambiguous rather than forcing a match.

A practical label set

DimensionLabelMeaning
PresenceNot presentEntity is absent from the answer
PresenceMentionEntity is named or unambiguously referenced
Source roleDisplayed sourceA link or citation points to the entity's page
SupportFull, partial, none, unclearCitation support for the associated claim
StanceDescriptiveFactual or neutral statement without selection guidance
StanceConditionally suitableFit is stated for a persona, constraint, or use case
StanceRecommendedAnswer advises choosing or considering the entity
StancePreferred or rankedAnswer places the entity above named alternatives
StanceWarningAnswer identifies a material risk or limitation
StanceExcludedAnswer advises against the entity for the stated task
StanceAbstentionAnswer declines to choose or says evidence is insufficient
ConditionExplicit or inferredWhether the use-case boundary is stated in the text

Keep source role and stance in separate columns. A displayed source can support any stance. A mention can occur without a citation. A recommendation can be conditional and still meaningful, provided the condition is retained.

Citation support is a separate judgment

The paper Evaluating Verifiability in Generative Search Engines defines citation recall as the share of verification-worthy statements fully supported by citations and citation precision as the share of citations that support their associated statements. Its reported values came from older versions of four systems and should not be treated as 2026 platform performance. The annotation concepts remain useful.

For each cited claim, ask whether the source provides full support, partial support, no support, or cannot be assessed. A product page that confirms an integration may fully support “Platform C does not list Connector Z.” It does not support “Platform C is a poor choice for every enterprise.” The recommendation adds a decision rule that may require more evidence.

Multiple citations can jointly support one claim. Record the union only after checking the contribution of each source. A pile of links should not be accepted as support merely because one of them might contain the fact.

Platform metrics have their own boundaries

Bing's AI Performance announcement defines total citations, average cited pages, sampled grounding queries, and page-level citation activity. It explicitly states that these measures do not indicate placement, ranking, authority, page importance, or the role of a page within an individual answer.

That product boundary should survive export into an internal dashboard. A field named recommendation_rate cannot be calculated from Bing citation counts alone. It requires captured answer text and a defined annotation process. Likewise, a grounding-query sample is not a complete list of user prompts.

OpenAI's web-search documentation describes citations in generated responses and requirements for making them visible and clickable in applications. That tells developers how attribution is represented in that product path. It does not assign commercial sentiment to the cited entity.

Build the annotation handbook

  1. Write the target question. Decide whether the study concerns discovery, source attribution, factual support, recommendation, or rejection.
  2. Define the observation. Capture prompt, answer, citations, platform, date, market, language, account state where relevant, and run identifier.
  3. Resolve entities. Map names and product variants to stable IDs while retaining ambiguous cases.
  4. Segment claims. Mark the smallest clause that expresses a proposition or decision about an entity.
  5. Label presence and source role. Record mention and displayed citation independently.
  6. Assess citation support. Use full, partial, none, or unclear and save the evidence span.
  7. Label stance. Choose descriptive, conditionally suitable, recommended, preferred, warning, excluded, or abstention.
  8. Capture conditions. Record persona, market, budget, feature, availability, risk, or other decision boundary.
  9. Double-code a sample. Have independent annotators label the same observations before full production.
  10. Resolve disagreements by rule. Update examples and definitions instead of silently choosing one annotator's judgment.
  11. Version the codebook. Recalculate or clearly separate results when definitions change.

The OpenAI evaluation best-practices guide recommends task-specific evaluation, clear criteria, and continuous evaluation. A GEO annotation handbook should follow the same discipline: labels need examples, edge cases, and a stable relationship to the business question.

Edge cases that need explicit rules

List without guidance. “Options include A, B, and C” is a mention for each entity. It is not automatically a recommendation. If the answer introduces the list as “recommended options,” label the recommendation at the list-introduction level and attach it to each included item, subject to any item-specific caveat.

Citation after a paragraph. A citation marker at the end may appear to support several sentences. Annotate its association according to the interface and wording; use unclear when scope cannot be determined.

Recommendation by elimination. If the answer rejects A and B and states that C meets the requirement, C may be the selected option even without the word “recommend.” The codebook should define whether this counts as recommendation or preferred-by-elimination.

Quoted recommendation. “Analyst Firm X recommends A” is a report about another source's stance. Record the attribution. Do not label it as the assistant's own recommendation unless the answer adopts that view.

Negative citation. A source may be cited to document a recall, missing feature, or restriction. The citation is real and the stance is negative or cautionary.

Mixed persona. “A suits freelancers, while B is better for regulated enterprises” contains two conditional recommendations. The persona conditions are part of the label.

A worked annotation example

Question: “Which inventory system should a 40-person UK medical-device distributor choose if it needs lot tracking and Xero integration?”

Answer excerpt: “HarborTrack supports lot tracking and lists a Xero connector, so it is the strongest fit among these options.[1][2] LedgerNest also connects to Xero but its public documentation does not show medical-device lot workflows.[3] Ask each vendor to confirm UK implementation and validation requirements.”

For HarborTrack, mark mention, recommendation, preferred/ranked, and two displayed sources. Split “supports lot tracking” and “lists a Xero connector” into claims and test each citation against its source. Record the condition as a UK medical-device distributor with those two requirements. The final validation sentence qualifies the recommendation.

For LedgerNest, mark mention and warning or limitation, not rejection, because the answer says public documentation is missing rather than asserting the capability is absent. Citation [3] can support what the public documentation shows; it cannot prove that the product lacks the workflow.

The example is fictional. Its purpose is to show how wording strength, missing evidence, and recommendation status remain separate.

Measure agreement before publishing a rate

Percent agreement is easy to understand but can look high when one label dominates. Add an agreement statistic suited to the label type, and publish the coding sample size, number of annotators, adjudication process, and codebook version. The choice of statistic should be documented rather than treated as a badge.

Review disagreement by category. If annotators disagree mainly between “descriptive” and “conditionally suitable,” the codebook may need a rule for phrases such as “works well for.” If they disagree on citation scope, the interface capture may be inadequate. Disagreement is diagnostic evidence about the measurement system.

Do not train an automated classifier on unstable labels. Stabilize the handbook and maintain a human-reviewed test set first. Then report classifier performance separately for each consequential class, especially recommendation and exclusion.

Calculate denominators that answer real questions

Mention rate is the share of eligible answer observations in which the entity is mentioned. Citation rate is the share with at least one displayed source from the entity's domain or defined source set. Recommendation rate is the share with a recommendation label. Positive-with-condition rate is the share with conditional suitability or recommendation and an explicit condition. Exclusion rate is the share with an exclusion label.

State what counts as eligible. If the question asks for accounting software and a brand sells only hardware, including that observation in the brand's recommendation denominator may be misleading. Preserve both panel-wide and relevant-opportunity denominators when useful.

Repeated runs are observations of one prompt-platform case, not independent user demand. Summarize within case before making population claims.

Failure modes

The most common failure is calling every source display positive. Another is using sentiment analysis on the full answer and assigning the result to every brand. A generally positive paragraph can contain one serious warning about a specific product.

Teams also infer recommendation from order, even when the interface or answer provides no ranking language. Position may be measured separately. Finally, rewritten summaries can destroy the evidence trail. Keep the raw captured answer, citation target, timestamp, and annotation span.

Annotation quality checklist

  • The business question determines the label set.
  • The unit is an entity-claim inside a captured observation.
  • Mention and citation are separate binary fields.
  • Citation support and answer stance are separate judgments.
  • Conditions and personas stay attached to recommendations.
  • Warnings, exclusions, and abstentions have distinct labels.
  • Ambiguity can be recorded without forced classification.
  • Independent annotators double-code a documented sample.
  • The codebook includes edge cases and version history.
  • Rates publish their denominator and opportunity rule.

Frequently asked questions

If a brand is first in a list, is it recommended?

Only if the answer or interface provides selection or ranking meaning. Record list position separately from recommendation stance.

Can a citation be negative visibility?

Y

es. A source can support a warning, limitation, correction, or exclusion. Citation presence alone has no positive or negative sign.

Should a conditional recommendation count?

Count it in its own class and retain the condition. Combining it with unconditional recommendation loses the part most useful to buyers.

Can an LLM annotate other LLM answers?

It can assist after the codebook is stable. Validate it against a human-reviewed set, monitor class-level errors, and route uncertain or consequential cases to people. Do not treat agreement with itself as accuracy.

Does a cited page necessarily shape the answer?

  1. Displayed attribution, textual support, and causal influence are different questions. Citation-absorption research proposes proxy measurements for textual contribution, but those scores do not expose a commercial model's internal causal process.

Source and method note

Sources were retrieved on September 23, 2026. Bing and OpenAI documentation describe their own citation products. The verifiability paper supplies support labels and historical evaluation concepts; its system results are not current performance claims. Citation-absorption measures remain research proxies. The taxonomy, fictional example, formulas, and workflow are editorial proposals. No live answer panel was annotated, no platform ranking mechanism is claimed, and no named human review was conducted.

Back to insightsMarkdown version