AI Visibility Tools: What They Measure and What They Cannot

On October 18, 2025, more than 85% of tracked brands lost visibility in ChatGPT on the same day. Profound measured it across millions of prompts. Average brand visibility fell 31%.

None of those brands did anything to cause it. ChatGPT changed how it handles brands as entities, and the answer set contracted for everyone at once.

Every AI visibility dashboard on the market showed that day as a decline. Most of them showed it as your decline.

Understand that gap before you put a platform on the budget. The tools measure something real. They report it with a precision the data does not support. Almost none of them show you the platform baseline that would tell you whether the drop was yours.

What does an AI visibility tool actually do?

It runs prompts, reads the answers, and counts how often your brand appears.

That is the whole mechanism. A tool sends a set of questions to ChatGPT, Perplexity, Google AI Mode and others. It parses each answer for brand names and cited sources. Then it rolls those observations into a score, a share of voice figure, or a rank.

The measurement is honest work. The rollup is where the trouble starts.

Why do two tools disagree about the same brand?

Because they ask different questions, a different number of times, in different sessions.

Four variables move the number, and no two vendors set them the same way:

  • The prompt set. Different questions produce different answer sets.
  • The run count. One run per prompt per day is a coin flip recorded as a fact.
  • The session state. A signed in session carries personalization into the result.
  • The parser. Logged out, ChatGPT renders citations as publisher labels with no links, so a tool reading domains gets different data than one reading labels.

Then there is the variance underneath all of it. SparkToro found fewer than one in a hundred answer pairs name the same list of brands. Kevin Indig found only 2.2% of citations survive three runs of the same prompt.

A dashboard reporting 14.3% visibility to one decimal place is reporting a sample as a measurement.

What is the private index problem?

You cannot reproduce the number, so you cannot defend it.

Call it the private index. The vendor picks the prompts, hides the run count, runs them from accounts you do not control, and hands you a figure. When your CEO asks why visibility dropped four points, the honest answer is that you do not know, and neither does the vendor.

Compare that to a number you generate yourself. Ten prompts you wrote. Five runs each. Three engines, logged out. Anyone on your team can rerun it and get the same range. That figure survives a board meeting.

Do the tools show you when the platform moves?

Rarely, and this is the single most expensive gap.

Platform-wide changes are common now. The October 2025 entity update cut the number of brands named per response from six or seven down to three or four. Then GPT-5.3 Instant became the default on March 4, 2026. Resoneo’s study of 27,000 responses found unique domains cited per response fell from 19 to 15. It never recovered.

Fewer brands named per answer means lower visibility for almost everyone. Your dashboard draws that as a cliff on your chart.

Movement across a model change is unattributable. If a tool cannot tell you which model version produced each observation, it cannot separate your performance from the platform’s weather.

Seven questions to ask on the demo

Ask these before pricing comes up. The answers sort the market in about ten minutes.

  1. How many runs per prompt, per engine, per collection? One is not a measurement.
  2. Logged in or logged out? A signed in session puts personalization into your data.
  3. Which engines, and what happens to my history when you add one? Adding an engine changes the denominator and voids your baseline.
  4. Do you record the model label on every observation? Without it, no movement is attributable.
  5. Can I export the raw observations? Not the chart. The runs.
  6. Do you publish confidence intervals or run-to-run variance? Almost nobody does. Ask anyway.
  7. What do you show me when the entire platform moves? Look for a category or peer baseline, not an alert.

If the answers are vague, you are buying a number you cannot reproduce.

When is a tool the right purchase?

When the work is continuous and the surface is large.

Buy one if you track many brands or many categories. Buy one if you need daily collection across dozens of prompts. Buy one if you are an agency reporting to clients every month. That volume is real labor, and software does it cheaper than a person.

Skip it if you have one brand, one category, and no baseline yet. You do not have a monitoring problem. You have a diagnosis problem, and a dashboard will not tell you why you are absent.

Most teams buy the dashboard first, watch a number move for two quarters, and never learn which of the four conditions they fail. Improving visibility runs in a fixed order. Crawler access, then a claim in the initial HTML, then a claim worth selecting, then presence on the sites the engines cite. A tool measures the symptom. It does not find the cause.

What does a fixed protocol give you that a tool does not?

A number you own, and a reason behind it.

Run ten prompts, five runs each, across three engines and you get 150 observations. You also get the two things a score hides. You get the list of competitors named in your place, and you get the domains those answers were built from. That list is a work order. A visibility score is a mood ring.

Freeze the protocol. Run it again in 90 days after you have shipped fixes with dates attached. Movement with no shipped fixes is noise, and now you can tell the difference.

Ask before you sign

Take one question into every vendor call. Show me the raw observations behind this number, with the model label on each one.

A vendor who can produce that is selling measurement. A vendor who cannot is selling a chart.

The audit gives you the baseline first. It runs 150 observations across ChatGPT, Google AI Mode and Perplexity, logged out and reproducible. It scores 30 fixed sections on whether machines can reach you, read you and select you. It hands back the competitors named in your place.

See what the audit covers.