ChatGPT’s crawler spent 34.82% of its fetches on pages that returned a 404, according to a December 2024 Vercel and MERJ analysis. Googlebot hit 404s on 8.22% of its fetches. A third of one engine’s crawl budget landing on dead pages is not a content problem.
That number explains why so much AI visibility work goes nowhere. Teams rewrite pages, publish more posts, and add markup. The engine never reads any of it.
Improving brand visibility in AI search follows a fixed order. Access first. Then readability. Then the claim itself. Then presence on the sites the engines already trust. Work it out of order and every later fix measures nothing.
What decides whether an AI engine names your brand?
Four conditions, and each one depends on the one before it.
- The engine can reach the page. Its retrieval crawler gets through robots.txt, the firewall, and the CDN.
- The engine can read the claim. Your primary claim sits in the initial HTML, not behind JavaScript.
- The claim is worth selecting. It says something a competitor could not say unchanged.
- Other sites say it too. The answer gets assembled from sources beyond your own pages.
The fourth condition carries more weight than most teams expect. Muck Rack’s May 2026 analysis found 84% of citations in AI answers point at earned media rather than the brand’s own pages. Ahrefs found branded web mentions correlate 0.664 with AI Overview presence, ahead of every link metric they tested.
Your website is necessary. It is rarely sufficient.
Call the failure the locked door. A team spends a quarter on conditions three and four while condition one quietly fails. Every sentence they improve sits behind a door the engine cannot open.
Step one: can the engine reach you?
Check your robots.txt against two classes of bot, because they do completely different jobs.
Retrieval bots fetch your page to answer a question and cite you. These include OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User. Blocking any of them removes you from that engine’s answers.
Training bots collect content to train future models. These include GPTBot, ClaudeBot, CCBot and Bytespider. Blocking them is a legitimate business decision.
Blocking training is a choice. Blocking retrieval is a defect. Most sites that block retrieval never decided to.
robots.txt is only half the check. Your CDN or firewall can block bots at a layer robots.txt never shows. Cloudflare began blocking AI crawlers by default on new domains on July 1, 2025. It has since split its controls into search, agent and training categories. On September 15, 2026, it added a setting to disallow training while staying in search. If your site went live on Cloudflare after July 2025, check which setting someone picked.
Then check the 404s. Dead URLs that redirect to a friendly “page not found” screen answer with a 200 status. Every dead link on your site looks valid to a crawler that already wastes a third of its fetches.
Step two: is your claim in the initial HTML?
Open your homepage, view the source, and search for your headline sentence. If it is not there, most answer engines never see it.
The same Vercel analysis found none of the major AI crawlers execute JavaScript. ChatGPT’s crawler fetches JavaScript files on 11.5% of its requests and runs none of them. Only Google’s and Apple’s crawlers render.
Hero claims, capability lists, tabbed content and carousels are the usual casualties. A defensible sentence that arrives after hydration scores zero for retrieval. The fix is usually cheap. Move the claim into static HTML.
Step three: does your claim do any work?
Run one test on every claim on your homepage. Could a direct competitor publish this exact sentence on their site without it becoming false?
If yes, the claim does no work. “AI-powered platform that streamlines your operations” is true of forty companies. An engine assembling an answer has no reason to pick you out of that set.
This is where most visibility problems actually live, and it is why publishing more rarely helps. Ten new posts built on an adoptable claim produce ten more pages that any competitor could have written. The engine still has no basis to choose.
Why ChatGPT recommends your competitors instead of you walks through the full test, with a method for scoring your claims against a competitor’s real sentences.
Step four: where do the answers actually come from?
Look at the sources the engines cite in your category, then get on them.
You collected those domains if you ran the count in how to track brand mentions in AI search. Rank them by how often they appear. That ranking is your work order.
Review platforms show up often. SE Ranking found 34.5% of AI Overviews on commercial queries cite at least one review platform. G2 holds 23.1% of those links and Capterra 17.8%.
A profile matters. Volume past a profile matters far less. Seer Interactive found brands with no verified profile were cited 1% of the time. Brands with as few as 1 to 13 reviews reached 53.5%. Kevin Indig’s study for G2 found review volume explains under 2% of citation variance across 500 categories.
Get listed. Stop chasing review counts.
In some categories the entire answer set rests on a single industry directory. No generic list of “sites AI cites” would have named it. Derive the list from your own runs.
What should you skip?
Three tactics show up on every AEO checklist and have failed when tested.
llms.txt. SE Ranking studied about 300,000 domains and found no measurable relationship between the file and citation frequency. Ahrefs found 97% of llms.txt files received zero requests in May 2026. If you already published one, leave it and stop maintaining it.
Schema for citations. Ahrefs added schema to 1,885 pages and AI Overview citations fell 4.6%. ChatGPT and AI Mode moved within noise. Organization markup with accurate sameAs links still earns its place for telling machines which company you are. More schema types do not buy citations.
Chasing reviews. Covered above. A profile is an eligibility gate. Volume is not a lever.
What does this cost you in Google?
Some fixes help both channels. Some trade one for the other. Know which before you ship.
The two channels reward different pages. A Search Engine Land study of 10 sites found 49 of the top 100 organic pages had zero LLM traffic. Traffic quality differs too. Kaiser and Schulze found organic search referrals convert about 13% better than ChatGPT referrals.
Rewriting a title that currently ranks in Google, to chase a broader AI answer, can cost you organic clicks. Say so before you change it.
Four fixes serve both channels at once:
- Original screenshots, diagrams and data in place of stock images
- A named author with a real bio and a resolvable author page
- Methodology shown on the page rather than asserted
- Pages that answer the question in the first sentence under each heading
These are the cheapest wins to get approved by someone who only measures organic.
Ask this before you hire anyone
Any agency proposing AI visibility work should answer one question before it quotes. Which of the four steps are we actually failing, and how do you know?
If the proposal starts at step three or four without checking one and two, you are paying to improve pages behind a locked door.
The audit answers that question with evidence. It checks crawler access and rendering. It scores whether your claims survive a competitor copying them. It runs 150 observations across ChatGPT, Google AI Mode and Perplexity. Then it hands you a fix list in the order that matters.
Scott King helps enterprise technology companies tell better stories to scale sales. Automation, GTM systems, and writing on AI search.
