Direct answer
A synthetic source loop occurs when AI-derived text is published, loses a clear connection to its inputs, and later appears as independent evidence in search, research, or another AI answer. Investigate the loop through source lineage, claim history, distinctive wording, publication dates, ownership, and the underlying evidence event. Do not rely on an AI-text detector to decide authorship or truth. Mark uncertain origin as unknown. Keep web citation recirculation separate from recursive model training: the first concerns apparent evidence on the web; the second concerns generated data entering later training sets. Provenance can expose derivation, but it cannot certify that a claim is correct.
How one unsupported sentence becomes a chorus
A company publishes a vague claim that its material “reduces energy use by up to 30%.” An AI writing tool turns the claim into a supplier profile. Several directories copy the profile, a comparison site summarizes the directories, and a new AI answer cites two of those pages. The answer now appears to have independent support even though every path returns to the same unattributed sentence.
The problem is not that software touched the text. A carefully sourced AI-assisted draft can be more traceable than hurried human copy. The problem is broken lineage. Readers can no longer see which page measured the result, what “up to” means, which process was compared, or whether any measurement occurred.
Volume then disguises dependence. Five domains look like five witnesses. In reality, they are five reports of one unsupported assertion.
Two loops must remain separate
The web evidence loop happens when generated or derivative pages circulate through retrieval and citation. Its risks include false corroboration, source laundering, stale facts, and the loss of qualifications. A content auditor can inspect public pages, dates, phrases, links, and ownership to investigate this loop.
The training recursion loop happens when model-generated data enter datasets used to train later generations of models. The Curse of Recursion, later published in Nature under a model-collapse title, studies this learning problem across theoretical and experimental settings. The authors describe loss of information about the tails of the original distribution when generated data recursively contaminate training. Their work does not show that a cited page entered the training data of any named commercial system.
Do not use model-collapse research as proof that a particular AI answer copied a particular web page. Likewise, finding a web citation loop does not reveal how the underlying model was trained. The mechanisms, evidence, and responsible actors differ.
“AI-generated” is not an evidence verdict
Origin and validity are different dimensions. A human can invent a claim. A model can accurately summarize a well-documented standard. An edited document can combine primary measurements, licensed material, and generated connective prose.
An audit should therefore record at least three judgments: origin confidence, derivation quality, and claim support. Origin confidence describes whether the creation path is documented. Derivation quality describes whether sources and transformations are traceable. Claim support describes whether appropriate evidence backs the proposition.
Text classifiers cannot reliably settle all three. A detector score is a model output shaped by language, length, genre, and training data. False positives can harm authors, especially writers using formulaic, translated, or second-language prose. Use detector output, if at all, as a weak triage signal that never becomes the published finding by itself.
A source-lineage model
| Record | Key fields | Question answered |
|---|---|---|
| Claim | Exact proposition, subject, scope, units, date | What is being asserted? |
| Evidence event | Test, observation, filing, interview, dataset, policy | What produced the underlying fact? |
| Source object | URL or file, publisher, version, locator, checksum | Where is the evidence recorded? |
| Transformation | Quote, summary, translation, calculation, generation, edit | How did one object become another? |
| Agent | Author, organization, software, reviewer, importer | Who or what performed the transformation? |
| Dependency edge | Derived from, quotes, republishes, syndicates, contradicts | How are two sources related? |
| Disclosure | Method, assistance, source list, limitations | What can a reader inspect? |
| Status | Primary, independent analysis, derivative, unknown, withdrawn | How should the source be counted? |
The W3C PROV primer describes entities, activities, agents, derivation, and revision. Those concepts fit a source-lineage graph even when a team stores it in a spreadsheet. The graph should point from a published sentence through each transformation to an evidence event.
The crucial node is the evidence event. Two pages can be independently written and still depend on one press release. Conversely, two analyses of the same public dataset may be distinct analytical events if their methods and calculations are genuinely independent. State which kind of independence is being claimed.
Provenance standards help with history, not truth
The C2PA specifications define a technical approach to content provenance and authenticity using manifests, assertions, claims, and cryptographic relationships for media assets. Such credentials can help describe origin and edits when the creation and distribution chain preserves them.
A valid provenance record does not mean the content is accurate. It can establish that a known actor signed a record or that an asset passed through declared edits. It does not verify a product claim, make a source independent, or prove that omitted history never existed. Credentials can also be absent because a tool or channel did not preserve them.
For web text, teams can borrow the principle even when C2PA is not the implementation: preserve source IDs, version history, transformation notes, timestamps, and hashes. Keep a human-readable source note so provenance is not available only to specialized software.
Investigate a suspected loop
- Freeze the claim. Record the exact sentence, subject, number, unit, qualification, page, and access time.
- Collect every cited source. Save the relevant passages and metadata rather than only a list of URLs.
- Find the earliest accessible occurrence. Use publication dates, archives, feeds, and document metadata while noting that the earliest found page may not be the true origin.
- Search distinctive phrase fragments. Rare wording, identical errors, unusual ordering, and shared examples can reveal derivation.
- Inspect outbound and inbound attribution. Follow citations, “according to” language, canonical tags, syndication notices, and republisher relationships.
- Resolve ownership. Sites on different domains may share a parent company, data feed, author, agency, or commercial relationship.
- Locate the evidence event. Ask what test, observation, filing, dataset, or interview produced the fact.
- Compare transformations. Record which conditions, caveats, and uncertainty were added or removed at each step.
- Classify independence. Separate independent evidence, independent analysis, derivative reporting, copied text, and unknown lineage.
- Publish the uncertainty. If origin cannot be established, say so. Do not convert resemblance into an authorship accusation.
This procedure investigates provenance. A separate subject-matter review determines whether the underlying claim is credible and applicable.
A worked fictional lineage
Suppose six pages say that a fictional coating “extends pump life by 2.4 times.” Page F is cited in an AI answer. F links to E, a comparison article. E names no study but closely paraphrases D, a distributor blog. D copies a paragraph from C, a supplier marketplace listing. C was populated from B, a manufacturer feed. B summarizes A, an internal case presentation covering one pump at one facility with no control.
The visible graph contains six URLs and one evidence event. The multiplier may accurately describe the observed case, but the public wording has lost the facility, pump, baseline, maintenance changes, and study design. It cannot support a general life-extension promise.
A defensible summary would identify the available origin, state that the public pages are derivative, describe the case scope, and avoid treating the six pages as corroboration. The audit should also leave room for unknown evidence that was not accessible.
Scaled content changes the risk profile
Automation can multiply a lineage error across categories, markets, and comparison pages before anyone verifies the first claim. Google's spam policies define scaled content abuse around generating many pages primarily to manipulate search rankings rather than help users, and the policy can apply whether automation, people, or a combination produced the pages.
That is a Google Search policy boundary, not a universal definition of low-quality AI text. The practical editorial control is broader: every scaled workflow should require source fields, duplication checks, claim-status gates, and sampling that inspects actual rendered pages. Generation speed should not exceed verification capacity.
Google's current AI optimization guide emphasizes unique, valuable, people-first content and warns against producing pages for every query variation mainly to manipulate generative responses. A source-lineage system helps editors see when a “new” page adds an evidence event and when it merely adds another layer of paraphrase.
Preserve original evidence and minority cases
Synthetic loops tend to flatten nuance. A precise range becomes a single number. A disputed finding becomes a consensus statement. Rare failure modes disappear from summaries because they are less common in the source text.
The model-collapse paper's concern about distribution tails is specific to recursive learning, yet it offers a useful editorial analogy: derivative summaries often preserve central claims and lose exceptions. The analogy should remain an analogy. Web editors can respond directly by keeping source excerpts, methods, limitations, counterevidence, and revision history accessible.
When a generated summary is useful, label its role. Link to original documents. State the selection and synthesis method. Preserve conflicting results instead of averaging them into a claim no source made.
Decide whether a source can support publication
| Source condition | Count as independent evidence? | Publication use |
|---|---|---|
| Primary record of the evidence event | Yes for that event | Use within its scope and method limits |
| Separate team repeats the measurement | Potentially | Compare methods, data, and conflicts |
| Independent analysis of the same dataset | Independent analysis, shared data | Attribute the common dataset |
| News report quoting the primary record | No new evidence event | Use for context; cite primary where possible |
| AI summary with complete source mapping | No by default | Use as navigation or synthesis, verify originals |
| Page with no sources and repeated wording | Unknown or derivative | Do not count as corroboration |
| Provenance credential but unsupported claim | Origin may be clearer | Truth still requires suitable evidence |
Independence is claim-specific. A page can independently verify company identity while deriving its performance statistic from the company.
Source-loop prevention checklist
- Atomic claims link to underlying evidence events.
- URLs that repeat one source share one lineage ID.
- Generation, translation, calculation, and editing are recorded as transformations.
- Source versions, locators, dates, and hashes are preserved where appropriate.
- AI-text detector scores never determine authorship or truth alone.
- Provenance credentials are not treated as factual certification.
- Scaled workflows cannot publish stale or unverified claims.
- Original methods, limitations, and counterevidence remain accessible.
- Unknown origin stays labeled unknown.
- Corrections propagate to every mapped derivative page.
Frequently asked questions
Does AI-assisted writing create a synthetic source?
Not automatically. The important questions are whether claims are supported, sources are traceable, transformations are disclosed appropriately, and the page adds useful analysis or evidence.
Can matching phrases prove that one page copied another?
They can support an investigation, especially when wording and errors are distinctive. They do not prove direction, authorship, or unauthorized copying by themselves. Dates and relationships may also be incomplete.
Does C2PA prove content is true?
- It can strengthen provenance and tamper-evidence claims within its trust model. The factual proposition still needs appropriate evidence.
Is model collapse already happening in a named search engine?
The cited research demonstrates a recursive-training risk in studied settings. It does not disclose the training composition or establish model collapse in a specific commercial search product.
How should derivative sources be cited?
Link the primary evidence when available and cite the derivative source for its genuinely distinct analysis or context. State shared lineage when multiple reports depend on the same event.
Source and method note
Sources were retrieved on September 23, 2026. W3C PROV and C2PA address provenance models and technical records; neither certifies factual truth. The model-collapse paper studies recursive training and is not evidence about a named platform's training data. Google documents its own spam and AI-search guidance. The lineage table, investigation method, fictional coating example, and independence rules are editorial proposals. No site was accused of synthetic authorship, no detector was run, and no named human review is claimed.