All insights
XINDAR INSIGHT

The Two Chinese AI Webs: Why App and Browser Results Need Separate Audits

Chinese generative search should be audited as a set of platform-interface systems, not as one market-wide answer engine.

Direct answer

Chinese generative search should be audited as a set of platform-interface systems, not as one market-wide answer engine. The app and browser version of the same brand can retrieve different sources, expose different citations, use different account and device context, and present different actions. A defensible audit therefore samples each interface separately, repeats each query, preserves answer and citation evidence, and reports results by platform, interface, session state, language, time, and intent.

A 2026 preprint examining four mainstream Chinese-language platforms found systematic App-Web source differences across its sample. It also found that only a minority of brands present in cited material were surfaced in answers, and that much contact information appearing in answers could not be matched to the contemporaneous crawled body text. Those findings establish measurement problems in that study, not universal performance rates for every platform, brand, or future date.

One brand name can hide several retrieval systems

A consumer may think of a search product as one service. An auditor must distinguish its browser site, mobile web, native app, mini program, logged-in state, model or mode selection, and any web-search toggle. These surfaces can have different release cycles, permissions, personalization, data partnerships, citation interfaces, and caching behavior.

The preprint What Do Chinese-Language Generative Search Engines Cite and Surface? studied Web and App interfaces across four platforms. Its controlled design covered eight interfaces, 614 queries, and three repetitions for each query-platform-interface combination. The authors report 214,119 raw records and a cleaned citation-level dataset of 160,860 records.

This scale is useful, but it does not turn the sample into a permanent census. Platform implementations can change rapidly. The query set, collection period, accounts, devices, location, and cleaning rules define what the estimates mean. Anyone using the paper should review the full current manuscript and supplementary material, not rely on the abstract alone.

Treat interface as an experimental stratum

An app result and a browser result should occupy separate rows even when they share the same platform logo. Pair them only after each is recorded independently. The minimum observation key is:

platform x interface x mode x account state x device x location x language x prompt variant x repetition x time

Changing one field creates a different condition. If the app is signed in and the browser is anonymous, an observed difference cannot be attributed to interface alone. If the app has location permission, local recommendations may differ for a legitimate reason. If the browser surface exposes citations while the app hides them behind a tap, visibility measurement also changes.

Use blocking or matched collection where possible. Run app and browser observations close in time, with equivalent account and location settings, randomize collection order, and repeat. Preserve the deviations that cannot be controlled.

Retrieval, citation, and answer exposure are separate datasets

The citation pool contains sources a response visibly attributes or that the collection method can extract. Answer exposure records brands, products, people, locations, claims, and contact details that the answer actually presents. The two sets overlap but are not identical. The preprint From Citation Selection to Citation Absorption proposes related measures for separating displayed selection from contribution to an answer; its proxies should not be mistaken for direct access to a model's internal causal trace.

LayerUnitExample measureMain limitation
ResponseOne generated answerWeb search activated, refusal, answer lengthInternal retrieval may be hidden
CitationOne cited URL or source cardDomain, rank, title, publication dateDisplayed citation may not support nearby claim
Entity exposureOne normalized entity mentionBrand selected into answerEntity resolution can merge or split names incorrectly
Contact exposureOne phone, address, account, or URLContact matched to cited body textData may come from snippets, databases, memory, or another source
Action exposureOne call, map, shop, or follow controlAction available above foldUI and permissions affect visibility

The study reports an overall brand-selection rate of 8.3 percent among brands in its citation pool. That denominator is essential: it is not the percentage of all Chinese brands visible in AI answers. It describes selective surfacing among brands identified in the cited-source records under the study design.

It also reports that 12.4 percent of retrieved sources containing contact information contributed contact information to answers. Again, this is a source-to-answer contribution measure under the authors' matching rules, not a business conversion rate.

Unmatched exposure is a provenance warning

Approximately 13 percent of brand exposures in the study could not be matched to the contemporaneous citation pool, and approximately 71 percent of contact-information exposures could not be matched to the crawled body text. An unmatched item is not automatically false. It may derive from a snippet, structured data, a page version not captured, another source, a platform database, model memory, or an entity-resolution error.

The correct label is "unmatched under this collection method." Calling every unmatched phone number a hallucination would overstate the evidence. Calling the same number verified because it appeared confidently would be equally unsafe.

Contact information deserves a dedicated validation path. Match the exposed value against the organization's official current page, authoritative business registry where relevant, and a direct ownership record. Record country code, formatting normalization, branch, service hours, verification date, and whether the number is still controlled by the entity. A stale phone number can cause direct harm even when the rest of the answer is accurate.

Freshness depends on query type

Among cited pages with publication dates, the study fit approximate half-lives of 39 days for high-timeliness queries and 68 days for low-timeliness queries. A fitted citation half-life describes decay in the observed distribution; it is not an expiry date for every page. Publication dates may be missing, inconsistent, or updated without substantive change.

Use freshness rules that match the fact. A restaurant phone number, product price, event date, regulation, and historical definition age at different rates. Record the validity time of the fact, not merely the publication date of the page. An older standards document can remain the primary source, while yesterday's price can already be stale.

For monitoring, create time cohorts and repeat the same query set. Report source-age distributions, new-source entry, removed-source persistence, and contact verification separately. Do not reward pages for date changes that lack material updates.

Source-set overlap needs a clear denominator

To compare App and Web, calculate overlap for each matched query and repetition. Jaccard similarity, intersection divided by union, is easy to explain. Also report directional containment: what fraction of App sources appear on Web, and what fraction of Web sources appear in App. A small app source set can be fully contained in a large web set while Jaccard remains modest.

Normalize URLs carefully. Remove known tracking parameters, resolve redirects, standardize scheme and host, and decide how to handle mobile subdomains, syndicated copies, and platform-native cards. Keep both normalized identity and raw URL so reviewers can reconstruct the choice.

Domain-level overlap and page-level overlap answer different questions. Two interfaces may cite the same domain but different articles. Report both when source identity matters. For entity exposure, use a separate normalization table with aliases, Chinese and Latin names, abbreviations, and parent-subsidiary relationships.

A rigorous Chinese AI audit protocol

  1. Freeze the research question. Decide whether the audit concerns source diversity, brand exposure, factual support, contact accuracy, or action visibility.
  2. Inventory interfaces. Record app version, operating system, browser, URL, model or mode, web-search setting, login state, and permissions.
  3. Build query cohorts. Include discovery, comparison, verification, local, contact, after-sales, risk, and current-event intents in natural Chinese.
  4. Create paraphrases. Preserve intent while varying wording and entity position.
  5. Schedule matched runs. Pair App and Web close in time, randomize order, and run at least several repetitions appropriate to the expected variance.
  6. Capture raw evidence. Save prompt, answer text, visible citations, expanded cards, screenshots, timestamps, and permitted network or referral data.
  7. Normalize sources and entities. Preserve raw values and document every rule.
  8. Code claim support. Check whether each citation supports the associated answer proposition.
  9. Verify sensitive facts. Recheck phone, address, price, regulation, medical, financial, and safety claims against current authoritative sources.
  10. Report strata and uncertainty. Keep App and Web results separate before any pooled view.

The protocol should include exclusion reasons: blocked collection, missing citation panel, login failure, CAPTCHA, truncated answer, or unresolvable source. Excluding failed runs without reporting them can bias the apparent rate of citation or search activation.

Use repetitions to reveal instability

Three repetitions, as used in the cited study, can reveal obvious variation but cannot estimate every rare outcome precisely. The needed count depends on the decision and expected variance. A high-stakes factual audit may use fewer prompts with deeper human verification; a visibility trend may use more prompts and repeated dates.

Report the number of distinct sources across runs, probability of any brand exposure, median and range of citation count, and consistency of key claims. Do not merge three different answers into one synthetic "best" answer and then call it the platform result.

Paraphrases matter because natural wording changes retrieval. Keep the original intent class stable, use a predeclared set, and avoid choosing only variants that favor the target brand. A balanced panel includes prompts that can surface competitors, neutral sources, negative evidence, and "none of the above" outcomes.

Validate automated coding

Large audits need automated extraction, but source cards, Chinese punctuation, aliases, and UI changes can break parsers. Build a labeled validation sample for each interface. Measure citation extraction precision and recall, entity resolution errors, contact matching accuracy, and claim-support agreement.

OpenAI's evaluation guidance recommends task-specific data, clear metrics, and continuous evaluation. It is general guidance from another platform, not validation of a Chinese search audit. The principle applies: test the evaluator against the actual failure modes of the collection system.

For claim support, blind reviewers to target-brand status where practical. Separate "citation is topically related" from "citation supports this exact proposition." Record disagreements and use an adjudication rule. An automated semantic similarity score can prioritize review but should not be treated as proof for sensitive claims.

A fictional paired audit

Suppose a medical-device distributor wants to know whether four platforms provide its correct after-sales phone number. The query panel includes official contact, repair booking, warranty service, local office, and "who should I call" variants. Each is run five times in app and browser over two collection days, with matched anonymous conditions where supported.

The audit records whether web search appeared active, which sources were displayed, the number shown, its location in the answer, and whether a call action was available. The number is then checked against the distributor's official service page and current switchboard record. A citation to the homepage is not counted as support if the captured page does not contain the number.

Results might show that one app consistently displays the current number from a platform business card while its browser interface cites an old distributor article. The conclusion is interface-specific: the app's contact result was accurate in the observed runs, while the browser source path was stale. The appropriate action is to update and deprecate the old article, strengthen the official contact record, and remeasure. It is not "Platform X knows the brand."

Governance and risk boundaries

Audit collection must respect platform terms, privacy, and security controls. Do not circumvent authentication, rate limits, or CAPTCHAs. Avoid storing personal account data unless necessary, consented, and protected. Contact details for public businesses can still become sensitive when linked to individuals or private channels.

NIST's AI Risk Management Framework offers a general governance structure for mapping context, measuring risk, and managing responses. It does not certify any audit method. Use it to assign owners, severity, escalation, correction verification, and retention policy.

The critical GEO survey also emphasizes platform and surface specificity, repeated measurements, and fidelity gaps. Its review is a preprint bounded to a defined literature window. Together with the Chinese-language empirical study, it supports a cautious operating principle: visibility is a distribution conditioned on interface and time.

App-Web audit checklist

  • Each platform-interface combination has a stable observation ID.
  • App version, OS, browser, mode, login, location, and permissions are recorded.
  • Query cohorts and paraphrases are fixed before target results are inspected.
  • Runs are paired in time and repeated in randomized order.
  • Raw answers, source cards, screenshots, and timestamps are retained.
  • URL, domain, entity, and contact normalization rules are documented.
  • Citation presence, claim support, entity exposure, and action exposure are separate.
  • Unmatched exposures are labeled without automatically calling them false.
  • Sensitive and time-varying facts are checked against current authoritative sources.
  • Pooled reports preserve App and Web strata and display sample size.

Frequently asked questions

Can an app and website be treated as the same platform?

Only at the corporate-brand level. For measurement, treat them as distinct interfaces until evidence shows aligned source and answer behavior for the target task.

Does an unmatched brand or phone number mean hallucination?

  1. It means the item could not be matched under the recorded citation and page-capture method. Investigate alternative provenance and verify the fact independently.

Are three repetitions enough?

They can reveal variability and support large-scale coverage, but adequacy depends on the outcome and precision needed. Report uncertainty and use more repetitions or dates for consequential conclusions.

Should results be pooled across Chinese AI platforms?

Only for a clearly defined market-level question, with platform and interface weights explained. Always preserve the disaggregated results because pooling can hide opposite behavior.

What is the highest-priority field to verify?

Prioritize facts whose error creates harm or misdirects action: contact details, price and availability, regulation, safety, medical, financial, and legal information. Priority still depends on the use case.

Source and method note

Sources were retrieved on September 24, 2026. The Chinese-language study supplies the reported design and descriptive estimates; the critical survey supplies a broader pipeline and measurement context; OpenAI and NIST supply general evaluation and governance guidance. Both research papers are preprints and may be revised. The audit key, protocol, overlap measures, validation design, and fictional distributor example are editorial proposals. No current platform ranking, client result, or named human review is claimed.

Back to insightsMarkdown version