Direct answer: An AI visibility audit should include a documented prompt baseline, answer and citation evidence, brand-accuracy review, competitor comparison, technical accessibility checks, entity consistency, content and source-gap analysis, risk findings, and a prioritized implementation roadmap. It should state what was sampled and avoid promising a permanent AI ranking.
The audit begins with scope, not screenshots
An audit is only interpretable when the reviewer knows what was tested. Before collecting answers, define the business unit, products, target countries, languages, buyer roles, competitors, platforms, and time window.
A US B2B software audit and a German industrial supplier audit should not use the same prompt set. Their category vocabulary, evidence, sources, competitors, and decision risks differ. Scope also determines which claims and pages can reasonably be reviewed.
The audit should list assumptions and exclusions. For example, it may exclude logged-in personalization, voice interfaces, non-English markets, or regulated claims that require specialist review.
1. A controlled prompt baseline
The prompt set should represent the buyer journey:
- category and problem discovery;
- approach and solution education;
- vendor and product comparison;
- implementation and integration;
- evidence, risk, policy, and compliance;
- branded description and reputation;
- follow-up questions that refine a shortlist.
Each prompt needs an ID, exact wording, intent, market, language, buyer role, funnel stage, and priority. Record the platform, interface, date, session method, and search or grounding state where observable.
The result is a baseline, not a universal census of everything an AI system may say.
2. Answer-level evidence
For every sampled answer, capture fields that a reviewer can validate:
| Evidence field | What to record |
|---|---|
| Brand presence | Mentioned, recommended, cited, or absent |
| Position | Order in a shortlist where one exists |
| Description | Category, capabilities, audience, market, differentiators |
| Accuracy | Correct, partly correct, incorrect, outdated, or unverifiable |
| Citation | URL, domain, claim supported, and source type |
| Competitors | Included brands and how they are framed |
| Risk | Hallucination, identity confusion, wrong price, policy, or availability |
| Evidence artifact | Compliant text capture, screenshot, or structured record |
Reports should distinguish observation from inference. “The brand was absent in 18 of 30 sampled prompts” is an observation. “The brand was absent because the site lacks authority” is a hypothesis that needs supporting technical and source evidence.
3. Citation and source analysis
An audit should identify which domains support answers in each prompt cluster. Classify official sources, independent authorities, industry media, directories, communities, competitors, and low-quality or outdated pages.
Then examine the selected passages. What question did the source answer? Was the claim direct? Did it include evidence, specifications, definitions, or comparison criteria? Was the page accessible and current?
The output should be a source-gap map, not a request to copy competitors. It should show which evidence assets and external relationships are missing.
4. Brand and entity accuracy
Review how answer systems identify:
- the company and preferred brand name;
- product and service categories;
- parent, subsidiary, and partner relationships;
- founders, experts, and authors where relevant;
- headquarters, markets, and contact information;
- product names, versions, and availability;
- credentials, certifications, and proof.
Compare the website with authoritative public profiles and documents. Flag conflicts with an owner and recommended canonical source.
5. Technical accessibility
GEO does not excuse weak technical SEO. The audit should review status codes, robots controls, indexability, canonical tags, sitemaps, internal linking, rendered text, mobile access, performance risks, and relevant structured data.
OpenAI publishes crawler documentation, and Google states that normal Search technical requirements apply to its AI features. Teams should review current platform documentation because crawler names and controls can change.
Structured data should be validated against visible content. Article schema, for example, should use an accurate headline, author, publication date, modification date, image, and publisher where available. Schema is descriptive evidence, not a guarantee of citation.
6. Content and answer readiness
Review priority pages for:
- a direct answer to the main question;
- descriptive headings and useful section structure;
- explicit definitions and category language;
- evidence attached to important claims;
- clear limitations and applicability;
- useful comparisons and decision criteria;
- first-hand expertise, examples, or methods;
- links to services, cases, policies, and source material;
- accurate author and update information;
- a clear next action for the reader.
Score page quality against the buyer question, not a generic word-count target.
7. Competitor comparison
Competitor analysis should show where another brand is selected, what claim earns attention, which sources support it, and whether your brand has an equivalent or stronger verified asset.
Useful categories include category ownership, answer coverage, source diversity, independent corroboration, product documentation, comparison criteria, expert authorship, market localization, and risk content.
Do not infer business performance from AI visibility alone. A frequently mentioned competitor may still have weak conversion, product fit, or customer satisfaction.
8. Risk review
High-priority risks include:
- confused identity or association with another company;
- fabricated product capabilities or certifications;
- outdated pricing, policies, addresses, or availability;
- unsupported health, finance, legal, security, or performance claims;
- negative claims sourced from obsolete or irrelevant material;
- inconsistent statements across markets and languages.
Each risk needs severity, evidence, owner, correction source, and monitoring cadence. Some issues require legal, security, quality, or regulatory review rather than a marketing rewrite.
9. A prioritized roadmap
The audit should end with decisions. A practical roadmap separates:
First 30 days
Fix crawl and index problems, resolve critical entity conflicts, publish corrections for high-risk claims, and establish the baseline measurement file.
Days 31 to 60
Improve priority service and product pages, publish missing evidence assets, strengthen internal links, and align authoritative profiles.
Days 61 to 90
Expand comparison and educational coverage, pursue relevant source relationships, localize priority markets, and repeat the controlled sample.
Every recommendation should include expected outcome, owner, effort, dependencies, evidence requirement, and measurement method.
Questions to ask an audit provider
- Which platforms, markets, languages, and interfaces will be sampled?
- Can we review the exact prompt set and scoring rubric?
- How are answer variability and personalization handled?
- Will the report include cited URLs and evidence artifacts?
- How are observations separated from causal hypotheses?
- Does the audit review technical access, entities, content, sources, and risk?
- Are competitor findings tied to specific prompts and sources?
- What implementation priorities and owners will be delivered?
- How will results be retested?
- Which outcomes are explicitly not guaranteed?
Common audit red flags
Be cautious when a report provides only favorable screenshots, hides the prompt set, combines all markets into one score, calls mentions “market share,” promises guaranteed rankings, or recommends mass content without a source and evidence diagnosis.
Another warning sign is a precise score with no rubric. Precision does not create validity. The score should be traceable to observable fields and a documented calculation.
Frequently asked questions
How long should an audit take?
A focused baseline can often be completed in one to two weeks, depending on markets, products, prompt volume, and review depth. Larger multilingual or regulated scopes require more time and specialist input.
Is an audit a one-time project?
The deep diagnostic can be periodic, but priority prompts and risk issues should be monitored continuously or at a defined cadence. Platform and market conditions change.
Should the audit include website analytics?
Yes, when available and appropriately governed. Search, content, conversion, and sales data help prioritize action. AI answer observations should not be interpreted as revenue impact without connecting them to downstream behavior.
What should the final deliverables be?
At minimum: scope, prompt inventory, baseline dataset, answer scorecard, citation map, entity findings, technical findings, content gap map, competitor findings, risk register, and 30/60/90-day roadmap.
Related Xindar services
- Request an AI Visibility and GEO Audit
- AI Visibility Monitoring
- AI Visibility Intelligence
- Contact Xindar
Sources
- Google Search Central, AI Features and Your Website. Accessed September 1, 2026.
- OpenAI Platform Documentation, Overview of OpenAI Crawlers. Accessed September 1, 2026.
- Google Search Central, SEO Starter Guide. Accessed September 1, 2026.
- Google Search Central, Article Structured Data. Accessed September 1, 2026.
- Schema.org, Article. Accessed September 1, 2026.
Method note: An AI visibility audit is a controlled sample of observable outputs. It cannot guarantee or fully enumerate every answer a platform may generate.