which ai engines can even read your site?
ChatGPT, Claude and Perplexity each send a named crawler to read your pages. One line in your robots.txt, or one firewall rule, and that engine never sees you, so it can never recommend you. In our July 2026 study of the Tranco top 1,000, 9.2% of the 653 reachable sites blocked OAI-SearchBot, the crawler that fetches sources for ChatGPT answers. We read your robots.txt the way those crawlers do and tell you exactly who gets in. No email.
free / no email / ~30s
crawler access is 1 gate of the full 164-check audit
In our scan of 1,000 sites, a significant share were blocking at least one major AI crawler — often without realising it.
Two kinds of crawler. Only one decides if you show up.
Every AI vendor documents its crawlers by name, and they are independent controls. We test each one against your live robots.txt using the vendors' own documented user-agent tokens, plus the wildcard and longest-match rules robots.txt actually uses.
- Search crawlers: the ones that make you visibleOAI-SearchBot (ChatGPT search), Claude-SearchBot (Claude search) and PerplexityBot build the live index those assistants answer from. Block one and you can be invisible in that engine's answers. This is the failure that costs customers.
- Training crawlers: a separate, legitimate choiceGPTBot, ClaudeBot, Google-Extended and CCBot collect pages for model training. Blocking them is a real opt-out that vendors document and respect, and it does not remove you from those engines' search answers. We report it, we don't judge it.
- llms.txt: optional, and we say soWe check whether you publish one. Google says plainly you do not need one to appear in AI search. It is a cheap experiment some non-Google tools read, not a requirement, and we will not scare you about it.
- The firewall behind the welcome matA robots.txt that allows a bot means nothing if your CDN or firewall 403s it at the door. We probe your page with real documented bot user-agents and report who gets turned away, with the honest caveat that we probe from a datacenter address.
If a bot you want is being turned away, the repair is in fixing AI crawler access and fixing robots.txt rules that block AI crawlers. For why any of this decides whether you appear in AI answers, start with what AEO is.
Once crawlers can reach your site, the next question is whether they find an llms.txt at your root — a short guide that tells AI tools which pages matter. The llms.txt generator builds one in under a minute.
AI crawler access is one of 12 signals that determine whether AI engines can find, read, and cite your site. Run a full GEO audit to see all of them together — llms.txt status, passage citability, answer-block presence, and entity clarity.
The honest version.
Which AI crawlers should I allow in robots.txt?
Allow the search-class crawlers if you want customers to find you through AI: OAI-SearchBot (ChatGPT search), Claude-SearchBot (Claude search) and PerplexityBot (Perplexity). Those are the ones that decide whether an engine can read and cite your pages in its answers. The training crawlers, GPTBot, ClaudeBot, Google-Extended and CCBot, only feed model training, so allowing or blocking them is a business decision about your content, not a visibility decision.
Does blocking GPTBot remove my site from ChatGPT?
No. GPTBot is OpenAI's training crawler; blocking it opts your content out of model training and nothing else. ChatGPT's search results come from a different crawler, OAI-SearchBot, which OpenAI documents as an independent control. The same split exists at Anthropic (ClaudeBot for training, Claude-SearchBot for search). Many sites copy a blanket block-everything robots.txt from a blog post and accidentally block the search crawlers too, which is the mistake this checker catches.
Do I need an llms.txt file to show up in AI answers?
No. Google's AI-search guidance says outright that you do not need new machine-readable files or AI text files to appear in generative AI search. An llms.txt is an optional, unofficial experiment: a short index of your key pages that a few non-Google AI tools read. It costs almost nothing to publish, so we report whether you have one, but no engine vendor requires it and skipping it does not hold you back with Google.
Access is the first gate. Then 164 graded checks.
Getting the crawlers through the door only matters if what they find is worth citing. The free whole-site scan runs 164 graded checks across SEO, answer eligibility, and AI readiness: full score, every gate failure, worst problems first, no email wall. Policy templates if you need to decide training vs search: AI bot policy. We make money on the keep-forever report and ongoing watch, not your inbox. Pricing. For AI-readiness beyond crawler access — answer signals, entity clarity, and citability — see the full AEO audit.