# The Content Settlement: Inside the Fight Over How AI Search Pays for What It Reads

> Perplexity is paying publishers 80% of a $42.5M subscription pool; CNN is suing it over 17,000 articles. The deal between AI and content is being written now.
Word count: ~2,400 · Data as of: September 2026 · Reading time: ~11 min

- Canonical: https://www.aixindar.com/news/the-content-settlement-inside-the-fight-over-how-ai-search-pays-for-what-it-reads
- Markdown: https://www.aixindar.com/news/the-content-settlement-inside-the-fight-over-how-ai-search-pays-for-what-it-reads.md
- Author: Daoyu Guan — https://www.aixindar.com/experts/daoyu-guan
- Published: 2026-09-03T03:44:45.939Z
- Last updated: 2026-09-03T03:44:46.022Z
- Evidence checked: Not separately recorded in CMS
- Editorial status: Published
- Corrections: No correction record supplied by CMS.

Two headlines from the same industry, months apart, define the question every content owner now faces. First: Perplexity launches a \*\*$5/month subscription that pays publishers 80% of revenue\*\*, seeded with a $42.5 million pool, distributed by formula — the first serious attempt to bill for AI's use of web content at the answer level. Second: **CNN sues Perplexity** for scraping 17,000 copyrighted articles, alleging its user agent "generally ignores robots.txt" while Cloudflare's investigation had already caught the engine's stealth crawlers masquerading as ordinary Chrome browsers.

One company, both headlines. That is not incoherence — that is a negotiation in progress. The settlement between AI search and the web's content layer is being written right now, in subscription revenue-share formulas, courtroom filings, and CDN defaults. This article maps the state of that negotiation: what's on the table, who's enforcing, what the reference prices are, and what a rational posture looks like while terms are still open.

## Why the Old Bargain Broke

The previous deal between search and content was simple: Google indexes your work, sends you the visitor, monetizes the visit indirectly through its own ads. That deal was never signed, but it was honored at scale for two decades — Google search referrals still outnumber ChatGPT's by roughly **500×** as of the Reuters Institute's 2026 analysis (1,300× including Discover).

The problem is the trajectory, not the level. As covered in our companion piece *The Vanishing Click*, roughly 68% of Google searches now end without a click, and Google's AI Mode — announced at more than one billion monthly users at I/O 2026 — runs a **93% zero-click rate** with referrals of just 1.6–2.5%. Chartbeat's panel of 2,500+ sites recorded Google organic news traffic down 33% globally between November 2024 and November 2025. The Reuters Institute's survey of 280 media leaders across 51 countries found publishers expecting search referrals to fall **43% over three years**, with one in five fearing declines above 75%.

AI platforms, meanwhile, crossed a billion monthly referrals to the open web in mid-2025 (SE Ranking) and keep growing. The content is needed more than ever. The traffic that paid for it is going away. Something has to replace the honor system.

## The Perplexity Experiment: Billing by Citation, Referral, and Task

Perplexity's Comet Plus program is the most detailed public attempt to date. The mechanics:

**The vehicle.** Comet, Perplexity's agentic browser, went from paid product (July 2025) to free globally (October 2025) to iOS and Android (March 2026), reaching roughly 3 million monthly active users with DAU up 320% year over year. Comet Plus, at $5/month, is the monetization layer on top.

**The split.** **80% of subscription revenue goes to publishers**, from a starting pool of $42.5 million. Twenty percent retained is roughly the inverse of what app stores charge — a deliberate signal to the content industry.

**The formula — and this is the genuinely novel part.** The pool is distributed across three signal types: (1) direct referral traffic from Comet, (2) citations appearing in AI answers, and (3) **content used by agents completing tasks**. That third category is the first commercial implementation of "agent traffic" billing — payment for content consumed by an AI acting on a user's behalf, where no human ever sees the page. Perplexity's scale context: about 45 million MAU, 1.2–1.5 billion monthly queries, and an ARR around $500 million as of April 2026, up 335% year over year. The average Perplexity answer cites five to eight sources — meaning the formula's denominator, and its unit economics, are already production-tested.

Whether 80% of a $5 subscription is *enough* is a separate question, and the honest answer is probably not yet: Reuters Institute respondents expect little from licensing — only 20% anticipate meaningful AI licensing revenue, 49% expect a little, 20% expect none. But the mechanism matters more than the current number. For the first time, citation counts, referral clicks, and agent usage are all metered, priced, and paid against. If that formula survives, it becomes the template every AI-content negotiation references.

## The Enforcement Frontier: Courts and the New Referee

The other half of the settlement is being written by people with subpoenas.

**The CNN lawsuit (2026)** alleges Perplexity scraped 17,000 articles and that Perplexity-User — the fetch agent that follows links users ask about — "generally ignores robots.txt." The technical backstory makes the complaint more than rhetoric. In August 2025, Cloudflare published evidence that Perplexity was operating **undeclared crawlers rotating user agents, IPs, and ASNs to evade no-crawl directives**, disguising themselves as ordinary Chrome browsers. Perplexity denies the characterization, but its own documentation concedes the design point at issue: Perplexity-User is officially described as "an agent, not a bot" — user-initiated fetches that robots.txt, a protocol written for crawlers, was never designed to govern.

This is the fault line the whole settlement runs through. Every AI vendor now runs **three distinct agent roles**: training crawlers (GPTBot, ClaudeBot), search-index crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot), and user-initiated fetch agents (ChatGPT-User, Perplexity-User, Claude-User). The middle tier largely honors robots.txt — block PerplexityBot and you lose Perplexity citations, cleanly. The third tier is a genuinely new thing: when a *user* asks an AI to read a page, is that crawling, or is it browsing? Legally and technically, the answer is unsettled, and the CNN case is where it may get settled.

**Cloudflare, meanwhile, has become the de facto referee.** Sitting in front of roughly a fifth of the web, it published the stealth-crawler evidence, and in August 2026 shipped **Bot Preference Sync** — tooling that automatically aligns robots.txt with a publisher's AI policy across three categories (search, agent, training), backed by BotBase, a directory of bot operators with declared behaviors. When the referee starts offering policy enforcement as a product feature, the handshake era of crawler norms is over.

## The Reference Price: What Content Is Worth

There is exactly one large-scale, functioning market price for AI access to web content: **Reddit's licensing deals, reported at roughly $140 million per year** from buyers including OpenAI and Google (CJR, 2025). The comparison is instructive in three ways.

First, it prices *structured, high-signal conversation data* — the exact corpus that is now the most-cited source class in AI search (Reddit is the #1 cited domain across five major platforms per Peec AI's 30 million-citation analysis). Second, it was negotiated from a position of strength: Reddit could credibly shut off access, because its content lives behind one domain with API gates. Publishers whose content is open by default have no such leverage — which is precisely why lawsuits and CDN-level enforcement are the fallback. Third, it establishes that AI platforms *will pay meaningful sums* when the alternative is losing the corpus. The $140M figure is the floor of the negotiation, not the ceiling of the market.

Against that backdrop, Perplexity's $42.5 million starting pool is an opening bid, and CNN's lawsuit is the counter. Where the final terms land — per-citation pricing, per-agent-task pricing, training-data licensing, or some bundle — will define content economics for the next decade.

## The Posture Matrix: What to Do While Terms Are Open

For a brand or publisher, the rational stance during an unsettled negotiation is to preserve optionality without forfeiting presence:

**1. Split your crawler policy by role, not by vendor.** The standard 2026 posture matrix:


|                            |                   |                       |                                |
| -------------------------- | ----------------- | --------------------- | ------------------------------ |
| Your goal                  | Training crawlers | Search-index crawlers | User-fetch agents              |
| Maximum AI visibility      | Allow             | Allow                 | Allow                          |
| Visible but not trained on | **Block**         | Allow                 | Allow                          |
| Search presence only       | Block             | Allow                 | Block                          |
| Fully closed               | Block             | Block                 | Block (network layer required) |


Block GPTBot and ClaudeBot while allowing OAI-SearchBot, Claude-SearchBot, and PerplexityBot, and you get citations without contributing free training data. Note the traps: legacy 2023-era robots.txt files that blanket-block every AI agent also erase search visibility; and user-fetch agents like Perplexity-User generally ignore robots.txt by design — controlling them requires network-layer rules (the vendors publish their IP ranges for exactly this purpose).

**2. Audit for accidental blocks.** Common findings: old Bingbot disallow rules silently removing you from Copilot's retrieval layer (Bing's index feeds both Copilot and parts of ChatGPT search); CDN challenge pages returning 403s to crawlers your robots.txt allows; regex user-agent filters written for GPTBot accidentally matching OAI-SearchBot.

**3. Meter your agent traffic now.** If agent-consumption billing becomes standard, historical usage data becomes negotiating leverage. Distinguish the three roles in your logs (user agents make this straightforward), and if you run an indexable content business, quantify how often AI agents fetch your material today. You cannot negotiate a rate card you never built the meter for.

**4. Keep the fastest indexing lanes open.** Bing's IndexNow protocol remains the quickest route into the retrieval layer that feeds Copilot — and AI engines disproportionately reward freshness (Growth Memo's 5.3-million-result study found freshness the only signal that predicted top-3 placement; author bios, schema, and word count were noise).

**5. Price your participation deliberately.** Joining Comet Plus's pool or a licensing program is a strategic choice, not a compliance task. The calculus differs by business: a subscription publisher with a strong brand has leverage to wait or litigate; a commerce site that *wants* AI agents to recommend and transact with it should welcome agent traffic and optimize for the task-completion signal — because for commerce, the agent *is* the customer.

## The Honest Bottom Line

The handshake era — search engines take content, send traffic, and everybody pretends the exchange is even — is ending, and nothing has replaced it yet. What exists instead is a live negotiation conducted in three venues at once: product (Comet Plus's 80/20 formula, with its unprecedented billing for agent task usage), litigation (CNN and the unresolved status of user-initiated fetch agents), and infrastructure (Cloudflare converting crawler policy from a text file into an enforceable platform feature).

The Reddit deal established that the platforms will pay real money when they must. The open question — currently worth far more than the $42.5 million in Perplexity's opening pool — is what a citation is worth when nobody clicks it, and what an agent's reading of your page is worth when no human was ever going to. Every content owner's leverage in that negotiation depends on two things they can control starting today: knowing exactly what the agents are taking, and being structurally worth citing enough that losing them would hurt the engines more than it hurts you.

---

## Sources

1. Perplexity (2025–2026). Comet and Comet Plus publisher program announcements; PerplexityBot / Perplexity-User official documentation.

1. CNN v. Perplexity (2026). Complaint filings regarding 17,000 articles and robots.txt practices.

1. Cloudflare (2025). Stealth crawler investigation, August 2025; Bot Preference Sync and BotBase announcements, August 2026.

1. Reuters Institute for the Study of Journalism (2026). *Journalism, Media, and Technology Trends and Predictions 2026.* 280 leaders, 51 countries; Chartbeat panel, 2,500+ sites.

1. SE Ranking (2025). AI platform referral volume analysis, mid-2025.

1. Columbia Journalism Review (2025). Reddit AI licensing revenue reporting (\~$140M/year).

1. Peec AI (2026). 30 million citation analysis. March 2026.

1. Growth Memo / Kevin Indig (2026). Freshness ranking-signal study. 5.3 million search results.

1. Semrush (2025). AI Mode zero-click and referral-rate analysis, September 2025.

1. OpenAI (2026). GPTBot, OAI-SearchBot, ChatGPT-User documentation. Anthropic (2026). ClaudeBot crawler documentation.

*Figures reflect research and filings published as of September 2026. Litigation outcomes and program terms are subject to change; treat deal-specific numbers as dated snapshots of a fast-moving negotiation.*

## Editorial references

- [Editorial policy](https://www.aixindar.com/editorial-policy)
- [Research methodology](https://www.aixindar.com/research-methodology)
- [Corrections policy](https://www.aixindar.com/corrections)
