What an AI Visibility Audit Should Include

Most AI visibility audits are a site crawl plus a screenshot of a ChatGPT answer.

A crawl tells you whether pages load. It says nothing about whether an engine has a reason to pick you. One screenshot is one run of a probabilistic system. That is a coin flip presented as evidence.

An audit earns its fee on three properties. The number is reproducible. The denominator never moves. And it refuses to recommend things that do not work.

What should an AI visibility audit measure?

Two questions, and most audits only attempt one.

Can a machine reach you, read you, and lift a passage from you? That is the technical half. Crawler access, server-rendered claims, crawl hygiene, entity markup, passage structure, and whether anything on the page is citable at all.

Does your positioning give an engine a reason to select you? That is the harder half. A claim a competitor could publish unchanged gives a retrieval system no basis to choose you out of the set.

Both failures produce the same silence. An audit that checks one and skips the other explains half your absence.

Why does the denominator matter?

Because a readiness percentage computed on a shifting list cannot be compared to anything.

Score the same fixed sections every time. Say an audit scores 24 items for one client and 31 for the next. Those percentages mean nothing side by side. They mean nothing against the same client in 90 days either.

Ask any provider how many sections they score and whether that list changes between engagements. A moving denominator is the easiest way to make progress appear.

The same rule governs the prompt set. Change the prompts, the run count, or the engine list between measurements and you have started a new baseline. Freeze the protocol and improve the website instead.

What does the observation protocol look like?

Ten prompts, five runs each, three engines, logged out. One hundred and fifty observations.

The details matter more than the total:

  • Nine of the ten prompts never name the company. Asking whether a named brand is good guarantees a mention and measures nothing.
  • One brand prompt stays in the set. If an engine cannot say correctly who you are and what you sell, no other finding matters.
  • Every run opens a new conversation. Context from the last run contaminates the next one.
  • Every run is logged out. A signed in session puts personalization into a client’s report.
  • Every run records the model label. Profound measured a 31% average visibility drop across tracked brands on a single ChatGPT change in October 2025. Movement across a model change is unattributable, and a report that hides it will sell you a platform shift as your own decline.

Presence comes back as a count, never a rank or a single score. SparkToro found fewer than one in a hundred answer pairs name the same brands. Two decimal places on that data is theater.

What belongs in the machine read?

Six checks, all publicly verifiable, none requiring access to your systems.

  1. Retrieval crawler access. Separate the bots that fetch pages for citations from the bots that train models. Check robots.txt and the CDN, because a firewall blocks at a layer robots.txt never shows.
  2. Server-rendered claims. Is the primary claim in the initial HTML. Most AI crawlers do not execute JavaScript.
  3. Crawl hygiene. Real 404 status codes, no soft 404s, no redirect chains, a sitemap whose URLs resolve.
  4. Entity markup. Organization and author markup with sameAs links that resolve, and the same company name everywhere.
  5. Passage structure. Question-shaped headings, answers directly beneath, no passage that collapses when lifted out of the page.
  6. Citable assets. Original data, sourced statistics, named quotes, real comparison tables, first-party media.

Each finding needs an owner and an estimate. Most belong to whoever controls the site and the CDN. That person rarely bought the audit. Name them early or the work stalls.

What should an audit refuse to recommend?

The refusal list is the fastest way to judge a provider.

llms.txt. SE Ranking studied about 300,000 domains and found no measurable relationship to citation frequency. Ahrefs found 97% of llms.txt files received zero requests in May 2026.

Schema for citations. Ahrefs added schema to 1,885 pages and AI Overview citations fell 4.6%. Organization markup still earns its place for entity disambiguation. More schema types buy nothing.

Review volume. A profile is an eligibility gate. Seer found brands with no verified profile were cited 1% of the time against 53.5% for brands with 1 to 13 reviews. Volume past that explains under 2% of citation variance.

Blocking Google-Extended to leave AI Overviews. It does nothing. AI Overviews retrieve from the live Googlebot index.

A provider with no refusals is selling a checklist. The checklist is cheap to produce and expensive to work through.

What should you receive?

Five things, and a report a colleague can read without having been on the call.

  • The verdict first, then the evidence behind it.
  • Your own sentences quoted, scored against competitors’ actual sentences.
  • Presence counts per prompt, per engine, with the engines named and any incomplete run disclosed.
  • A gap list and a fix list, kept separate. Gaps need a decision from a person. Fixes need an owner and an afternoon. Run them together and the fixes stall waiting on the decision.
  • A saved baseline file, so the 90-day re-measure compares against something real rather than a memory.

A method note closes it. What was read, on what date, which engines under what protocol, and what the audit cannot tell you. Server logs, GA4 and ad platform data sit outside a public audit, and any provider claiming otherwise is guessing on your behalf.

Five questions to ask before you buy

  1. How many sections do you score, and does that list change between clients?
  2. How many prompts, how many runs, which engines, logged in or out?
  3. Can I reproduce your number myself?
  4. What will you refuse to recommend?
  5. Do I get the raw observations and a baseline file I keep?

A provider who answers all five in one call is selling measurement. A provider who dodges three of them is selling a report.

My audit answers both halves. You get 150 observations across ChatGPT, Google AI Mode and Perplexity. You get 30 fixed sections, scored the same way for every client. You get your own claims scored against quoted competitor sentences. You get a gap list and a fix list kept separate, plus a baseline file for the 90-day re-measure.

Scope, terms and current pricing are on the audit page. See what the audit covers.