# Visibility Without Visits: What a Preregistered AI Search Field Experiment Proves

> A 2026 preregistered field experiment provides causal evidence that changing access to Google's AI search features changed external click behavior in its study population.

- Canonical: https://www.aixindar.com/news/visibility-without-visits-what-a-preregistered-ai-search-field-experiment-proves
- Markdown: https://www.aixindar.com/news/visibility-without-visits-what-a-preregistered-ai-search-field-experiment-proves.md
- Author: Daoyu Guan — https://www.aixindar.com/experts/daoyu-guan
- Published: 2026-09-24T10:07:48.230Z
- Last updated: 2026-09-24T10:07:48.341Z
- Evidence checked: Not separately recorded in CMS
- Editorial status: Published
- Corrections: No correction record supplied by CMS.

## Direct answer

A 2026 preregistered field experiment provides causal evidence that changing access to Google's AI search features changed external click behavior in its study population. Assignment to an AI Mode-only condition reduced click-through rate by 18.8 percentage points relative to current Google Search, while exposure to the no-AI intervention increased it by an estimated 8.8 percentage points. The study also reports worse user-experience and trust outcomes under forced AI Mode. These estimates do not prove that every AI summary reduces traffic, that all publishers lose the same amount, or that citation has no value.

The experiment is unusually informative because participants used Google in everyday browsing and were randomly assigned to interface conditions. It is also bounded: United States Chrome users, seven treatment days, a forced AI Mode condition, low baseline AI Mode use, and imperfect suppression of AI Overviews in the no-AI arm. Publishers should use it to revise measurement models, not to convert one result into a universal revenue forecast.

## Why this study deserves close reading

Most claims about AI search and publisher traffic come from trend lines, analytics panels, or comparisons between search pages with and without summaries. Those observations can be valuable, but the interface appears in response to query type and other factors that also affect clicking. People asking a short navigational question are not directly comparable to people asking a complex explanatory question.

The preprint [AI in Search Reduces Publisher Referrals Without Improving User Experience](https://arxiv.org/html/2608.18352) reports a preregistered randomized field experiment. Random assignment makes the groups comparable in expectation, allowing the authors to estimate the effect of the assigned experience within the study. Preregistration reduces the opportunity to choose hypotheses after seeing results, though it does not eliminate every analytical judgment.

This is a preprint retrieved on September 24, 2026. Readers should check its current review status, revisions, code, and replication materials before using it for a consequential decision. The paper's design and limitations matter as much as its headline.

## The intervention had three conditions

Participants first completed a three-day baseline period. They were then assigned for seven days to one of three conditions through a browser extension.


| Condition      | Interface treatment                                      | Primary comparison meaning                                          |
| -------------- | -------------------------------------------------------- | ------------------------------------------------------------------- |
| Current Search | Google Search was left unchanged                         | Observed product experience during study period                     |
| No AI Search   | AI Overviews were hidden and AI Mode searches redirected | Effect of exposure to suppressed AI features, subject to compliance |
| AI Mode Search | Searches were redirected to AI Mode                      | Effect of assignment to a forced AI Mode experience                 |


The analysis included 1,100 participants who made at least one search during the treatment period; 956 completed the post-experiment survey. Recruitment was limited to US adults who primarily used Chrome and Google Search, with participants recruited through Prolific and a Northeastern work-study program. The sample skewed younger and highly educated relative to the general population.

These details define the estimand. The result is not "the effect of AI on humanity." It is the effect of these interface assignments on measured behavior and perceptions among eligible participants during this short period.

## The main click result is a percentage-point change

Assignment to AI Mode reduced click-through rate by 18.8 percentage points relative to Current Search, with a reported 95 percent confidence interval from -22.2 to -15.3 points. Exposure to the No AI Search intervention increased click-through by an estimated 8.8 percentage points, with a reported interval from 2.3 to 15.3 points.

A percentage-point difference is not a percentage change. If a baseline rate were 40 percent, an 18.8-point decline would produce 21.2 percent, a relative decline of 47 percent. The actual interpretation should use the paper's outcome definition and reported baseline rather than inventing a convenient denominator.

The two effects also use different estimators. AI Mode routing succeeded for 94.7 percent of searches, and the paper discusses the assignment effect as intent-to-treat. The extension initially hid most detected AI Overviews in the no-AI condition, but a Google HTML change broke the detection logic; only 51.1 percent were hidden overall. The authors therefore report a local average treatment effect for exposure to No AI. It should not be presented as the simple effect of pressing a universal "remove AI" switch.

## The no-AI compliance failure is substantive

Interface experiments on live platforms can fail because the platform changes underneath them. In this study, suppression of AI Overviews declined from high initial coverage to zero after a markup change. That is not a minor technical footnote. It changes the population for which the no-AI estimate is interpreted and introduces assumptions needed for the local effect.

The paper explicitly discusses the exclusion restriction: assignment should affect the outcome through actual AI Overview exposure rather than another path. If the broken interface or extension behavior altered experience independently, that assumption could be weakened. The authors report both intent-to-treat and local estimates, which is more informative than hiding the failure.

For GEO teams, the lesson is methodological. A browser automation report that says a selector ran successfully is not evidence that the intended surface was removed or displayed. Capture the rendered interface, log intervention compliance per session, and preserve failure dates. Product markup is part of the experiment.

## User experience did not offset the referral decline in this test

The study's title says publisher referrals fell without an improvement in user experience. It reports that forced AI Mode reduced trust and several experience measures relative to Current Search, and that users expressed preferences against AI Mode in news comparisons. It also reports more substitution toward competing search engines under AI Mode assignment.

These results challenge a simple story in which fewer clicks necessarily reflect a better completed answer. They do not establish that every no-click session is bad. A user can obtain a correct fact without visiting a source, and some tasks are completed efficiently inside a result. The experiment shows that, for the tested forced experience and period, the measured experience gains did not compensate for the click loss in the authors' outcomes.

The distinction matters for publisher arguments. "Clicks fell" is an ecosystem and business concern. "Users were harmed" requires user evidence. This study offers such evidence for its conditions, while leaving other interfaces, adoption paths, and longer-term adaptation open.

## What the experiment does and does not identify


| Claim                                                         | Support status                     | Reason                                                       |
| ------------------------------------------------------------- | ---------------------------------- | ------------------------------------------------------------ |
| Forced AI Mode changed external click-through in the study    | Supported causally within design   | Random assignment and high routing compliance                |
| Suppressing AI features increased click-through for compliers | Supported under stated assumptions | Local effect with incomplete suppression                     |
| Every AI Overview causes the same click loss                  | Not supported                      | Effects can vary by query, layout, source, and user          |
| A citation has no business value                              | Not tested                         | Citation exposure and downstream value were not fully valued |
| Publishers lost a specific amount of revenue                  | Not established                    | Clicks were measured, not publisher revenue or profit        |
| Long-term organic adoption will match forced use              | Not established                    | Treatment lasted seven days and baseline AI Mode use was low |


The paper also reports domain-specific click outcomes, including reductions in the fraction of users clicking news, Reddit, and Wikipedia in the AI Mode condition. Those results concern the specified classifications and user fraction. They should not be converted into a forecast for a particular publisher without its query mix, audience, position, monetization, and substitute channels.

## Compare randomized and observational evidence correctly

Pew Research Center's [observational analysis of Google searches](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) examined browsing data from 900 US adults and separately collected the corresponding result pages. In the collected visits, users clicked a traditional result on 8 percent of pages with an AI summary and 15 percent without one; they clicked a cited source in the summary on 1 percent of visits with a summary. The analysis also reported more session endings on summary pages.

Pew's study describes a large observed association for its March 2025 browsing sample and April result capture. It cannot by itself isolate the summary's causal effect because queries that trigger summaries may differ from those that do not. The 2026 field experiment addresses causality through random assignment but has its own compliance, sample, duration, and treatment limits.

The studies are complementary. Observational work shows what happened across naturally occurring queries and interfaces. Randomized work estimates what changed under an intervention. Agreement in direction strengthens concern, but the numerical estimates should not be pooled as if they measure the same object.

## Citation and referral are separate economic events

A source can contribute to an answer without receiving a visit. That contribution may still have value: recognition, trust, branded demand, licensing evidence, lead influence, or correction reach. It can also have little measurable value, or even substitute for the publisher's own page. The outcome depends on the source, query, interface, and business model.

Cloudflare's discussion of the [agentic Internet and publisher economics](https://blog.cloudflare.com/agentic-internet-bot-report/) frames the underlying tension between machine use of content and business models funded by direct relationships. Cloudflare is an infrastructure vendor with its own product and policy position; its article is not neutral causal evidence. It is useful for identifying operational questions about crawler purpose, attribution, and compensation.

A publisher should maintain separate ledgers for answer presence, citation, referral, engaged visit, subscription, lead, sale, and retained revenue. Connecting those events may require privacy-safe attribution and experiments. A citation count alone cannot price the contribution.

## A measurement model for visibility without visits

Start with the query opportunity. Record whether an AI feature appeared, whether the publisher's source was retrieved or cited, whether the brand or claim entered the answer, whether a link was visible, whether the user visited, and what happened after the visit. Each stage needs its own denominator.

For example, referral rate among cited answers differs from referral rate among all monitored queries. The first can improve while total referrals decline if the site is cited less often. A conversion rate among AI referrals can rise while revenue falls if traffic volume contracts. Report counts and rates together.

Add value outside clicks only when it can be observed. Branded-search lift, survey recall, assisted conversions, and sales-call source mentions can be measured with controls, but each has attribution limits. Do not assign a fictional monetary value to an answer mention merely to make a dashboard complete.

## How publishers should respond

First, segment pages by economic role. A reference page may succeed by shaping accurate understanding even with few clicks. A calculator may need a visit. A subscription investigation requires enough unresolved value to justify opening the article. A local service page needs an action route.

Second, make citations useful entry points. The title and opening should tell the reader what the page uniquely contains: primary data, full method, interactive evidence, local detail, or a consequential caveat. Do not withhold the direct answer, but provide depth that a short synthesis cannot fully replace.

Third, grow direct relationships ethically. Newsletters, RSS, accounts, memberships, preferred-source controls, and events reduce dependence on one referral surface. The offer should follow demonstrated value and use clear consent.

Fourth, run page-level experiments. Compare formats, titles, unique assets, and calls to action among similar query classes. Search Console, analytics, referral parameters, and server logs can show parts of the path. [OpenAI's evaluation guidance](https://developers.openai.com/api/docs/guides/evaluation-best-practices) also reinforces defining objectives, datasets, metrics, and continuous evaluation when testing AI systems; it does not prescribe a publisher traffic experiment.

## Design a replication before making a forecast

A publisher consortium could preregister a study across query classes, markets, devices, and interface conditions. Define the primary outcome, treatment, unit of randomization, minimum detectable effect, exclusion rules, and compliance test before collection. Capture screenshots and source lists while respecting platform terms and participant privacy.

Measure both user and publisher outcomes: task success, accuracy, trust, time, external clicks, source diversity, engagement, subscription, and correction behavior. Include a longer period to test adaptation. Stratify by informational, navigational, local, shopping, and high-stakes queries rather than averaging them into one result.

NIST's [AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) provides a general structure for governing and measuring AI risk. It does not validate this proposed experiment. Its emphasis on context, measurement, management, and documentation is helpful when interface research affects users and publishers.

## Field-experiment reading checklist

- Confirm publication status, preregistration, data, code, and revision date.
- Identify the sampled population, recruitment route, market, browser, and period.
- Write the treatment exactly as implemented.
- Separate assignment effects from exposure effects.
- Inspect compliance, attrition, missing data, and interface failures.
- Read outcomes in their original units and denominators.
- Distinguish primary from secondary and exploratory results.
- Preserve confidence intervals rather than only point estimates.
- Do not infer revenue, long-term adaptation, or universal query effects without data.
- Compare observational evidence without treating it as the same estimand.

## Frequently asked questions

### Does the experiment prove AI Overviews always reduce publisher clicks?

1. The no-AI intervention supports a causal increase for compliers in this study, but suppression was incomplete and effects can vary across queries, users, layouts, and time.

### Why is the AI Mode result more straightforward?

Routing compliance was high, so the assignment contrast is easier to interpret. It still represents forced, nearly universal use during a short study, not ordinary voluntary adoption.

### Can a publisher apply the 18.8-point estimate to its traffic?

1. Its audience, query mix, baseline click rate, citation pattern, and interface exposure may differ. Use the estimate as evidence of a possible causal mechanism and measure the publisher's own outcomes.

### Does a no-click answer have no value?

Not necessarily. It may satisfy a user and may create attribution or brand value. Those benefits require separate evidence; they should not be assumed from citation alone.

### What is the most important limitation?

There is no single most important limitation. For AI Mode, forced adoption and short duration constrain generalization. For the no-AI arm, declining suppression compliance changes interpretation. The US Chrome sample also limits population reach.

## Source and method note

Sources were retrieved on September 24, 2026. The primary experiment is a 2026 preprint and supplies the design and reported estimates; Pew supplies observational browsing evidence; Cloudflare discusses publisher economics from a vendor perspective; OpenAI and NIST supply general evaluation and risk-management guidance. Exact claims were kept within the cited study conditions. The measurement model, response framework, and replication design are editorial analysis. No revenue loss, universal platform effect, client outcome, or named human review is claimed.

## Editorial references

- [Editorial policy](https://www.aixindar.com/editorial-policy)
- [Research methodology](https://www.aixindar.com/research-methodology)
- [Corrections policy](https://www.aixindar.com/corrections)
