Run the same prompt three times. Only 2.2% of the citations survive all three runs, according to Kevin Indig’s 2026 prompt tracking analysis on Search Engine Land. The answer changed. Your brand did not.
That gap is the whole problem with tracking brand mentions in AI search. Your CEO asks whether ChatGPT recommends the company. You check once, see nothing, and check again an hour later. This time you appear. Neither result tells you anything on its own.
You can track this. You track it as a count from a fixed protocol, not as a score from a dashboard.
Can you actually track brand mentions in AI search?
Yes. You record presence as a count out of a fixed number of runs, per prompt, per engine.
The unit looks like this. Prompt 3, ChatGPT, present in 2 of 5 runs. Prompt 3, Perplexity, present in 0 of 5. That is a measurement you can repeat in 90 days and defend to a CFO.
What you cannot do is report one number. A single score hides the variance that produced it, and the variance is the finding.
Why does the same prompt return different brands every time?
Because the systems are probabilistic, and the answer set moves on every run.
SparkToro tested this across thousands of prompts and found that fewer than one in a hundred answer pairs name the same list of brands. Same question, same engine, different companies.
One check tells you nothing. Five checks tell you a rate. That is the entire argument for run counts.
What makes a tracking protocol worth anything?
Six things, and all six have to stay frozen between readings.
- A fixed prompt set. Ten prompts. Four category level, three problem level, two comparison, one brand level.
- Nine of the ten never name your company. Asking whether a named company is good guarantees a mention and measures nothing.
- A fixed run count. Five runs per prompt per engine. That gives 150 observations across three engines.
- A fixed engine list. ChatGPT, Google AI Mode, Perplexity. All three answer and cite while logged out.
- Logged out, every run. A signed in session carries memory and personalization into your measurement.
- A new conversation every run. Context from the previous run contaminates the next one.
Here is the trap that eats most tracking efforts. Call it the moving denominator. You add an engine in March, drop two weak prompts in April, and sign in during May because the quota ran out. Now the number moves and you cannot say why. Every change to the protocol voids the baseline you were building.
Freeze the protocol. Improve your website instead.
Which ten prompts should you use?
Write the questions your buyer asks before they know your name.
Four category prompts come first. “Best tools for X.” “Who are the leading vendors in Y.” These test whether you exist in the answer set at all.
Three problem prompts come next. Describe the pain in the buyer’s words, with no product category attached. “How do I stop Z from happening.” Buyers ask this way more often than they ask for vendors.
Two comparison prompts follow. Name two competitors, not yourself, and see who gets added to the pair.
One brand prompt closes the set. Ask who your company is and what it sells. If an engine cannot answer that correctly, nothing else on the list matters.
Write at least one prompt in your own positioning language. Take the claim on your homepage and turn it into the question a buyer would ask. When the engine answers it with four other companies, you have found the sharpest problem you own.
How do you run the count yourself this week?
Block an afternoon. You need a browser, a spreadsheet, and the discipline to not sign in.
- Write ten prompts in the language your buyer uses. Take your homepage claim and turn it into the question a buyer would ask.
- Open a temporary chat. Run prompt one. Record the result.
- Close the conversation. Open a new one. Run prompt one again, four more times.
- Move to prompt two. Repeat through all ten.
- Finish one engine before you touch the next.
- Record the model version each engine reports that day.
Record four fields per run, and nothing else.
| Field | What you record |
|---|---|
| Present | Yes or no. Did the answer name your brand |
| Vendors | Every company named, in order |
| Domains | Every source cited |
| Model | The version label the engine reported |
The vendor column matters more than your own presence. When four competitors show up in your place across 15 runs, you are looking at the answer set your buyer sees. That list is your work order.
What does a count tell you that a score does not?
It tells you where you are absent, which is the only part you can act on.
A score of 14% tells you to try harder. A count tells you where. You appear in 4 of 5 runs on the brand prompt. You appear in 1 of 5 on one problem prompt, and 0 of 5 on all four category prompts. Those are three different problems with three different owners.
Zero presence on category prompts with strong presence on your brand name means the engines know you exist and do not consider you an answer. That is a positioning problem, and no amount of publishing fixes it.
Zero presence everywhere, including your own brand name, usually means the engines cannot reach your pages at all. That is a technical problem, and it gets fixed in an afternoon.
What breaks the number?
Platform changes break it, and they break it for everyone at once.
On October 18, 2025, ChatGPT changed how it handles brands as entities. Profound measured the result across millions of prompts. Average brand visibility fell 31%. More than 85% of tracked brands declined. The number of brands named per response dropped from six or seven to three or four.
Nothing those brands did caused it.
It happened again on March 4, 2026, when GPT-5.3 Instant became the default. Resoneo’s 27,000 response study found unique domains cited per response fell from 19 to 15. The figure never recovered over the following 14 weeks.
This is why you record the model label on every run. Movement across a model change is unattributable. Any tool that shows you a cliff without showing you the platform baseline is handing you someone else’s weather as your own forecast.
What do you do with the result?
Work the list in order. Access first, then readability, then the claim, then third party presence.
Check whether retrieval crawlers can reach your pages before you touch a single sentence. Check whether your primary claim sits in the initial HTML rather than arriving after JavaScript runs. Then check whether the claim would survive a competitor copying it word for word.
Then look at the domains you collected. Most answers get built from sources you do not own. The shortest path to presence usually runs through the sites already cited in your category.
Set a date 90 days out. Same ten prompts, same five runs, same three engines. Record what you shipped and when, because movement with no shipped fixes is noise.
Ask before you buy a dashboard
Any vendor selling you AI visibility tracking should answer four questions in the first ten minutes. How many runs per prompt? Logged in or logged out? Can I export the raw observations? And what do you show me when the entire platform moves?
If the answers are vague, you are buying a number you cannot reproduce.
The full protocol takes about 75 minutes of wall clock for 150 observations. Add the setup. Add the patience to keep the prompt set frozen for a year.
The audit runs all 150 for you. It scores 30 fixed sections on whether machines can reach you, read you and select you. It saves the baseline, so the 90 day re-measure compares against something real.
Scott King helps enterprise technology companies tell better stories to scale sales. Automation, GTM systems, and writing on AI search.