ToolsCrawler permissions

Check your AI crawler permissions.

Check robots.txt permissions for named AI crawlers and review available scanner probes. Missing measurements stay unknown; access alone does not establish that an AI system retrieved or cited your page.

Crawler permissionsMeasured rules and available probes

Free category scores and selected findings. Full findings and export files: $10 once.

free / no email / ~30s

sample readout run yours above
search crawlers: reported robots.txt permissions
oai-searchbotOpenAI, powers ChatGPT search✗ blocked in robots.txt
claude-searchbotAnthropic, powers Claude search✓ allowed
perplexitybotPerplexity answers✗ blocked in robots.txt
Training and content-use policies
gptbotOpenAI model training✓ allowed
claudebotAnthropic model training✓ allowed
google-extendedGemini training and grounding; not Google Search✓ allowed
ccbotCommon Crawl✓ allowed
the rest of the door
llms.txt⚠ not published (optional, not a Google requirement)
firewall✓ no user-agent-level block found
sample: two search crawlers restricted by robots.txtreview access
This is a sample so you can see the shape of the answer. Run your own URL above for the real one.

crawler access is 1 gate of the full 157-check audit

In our scan of 1,000 sites, a significant share were blocking at least one major AI crawler, often without realising it.

what this checker measures

Separate search access from training permissions.

ChatGPT, Claude and Perplexity each send a named crawler to read your pages. Robots.txt and firewall rules can restrict their access. Access alone does not guarantee indexing, citations or recommendations. In our July 2026 study of the Tranco top 1,000, 9.2% of the 653 reachable sites blocked OAI-SearchBot, the crawler that fetches sources for ChatGPT answers. Check the robots.txt permissions and scanner probes recorded for your site. A missing measurement is shown as unknown. No email.

Every AI vendor documents its crawlers by name, and they are independent controls. This checker separates the rules your robots.txt declares from requests made by our scanner. Results describe the path and time checked. Missing or unavailable evidence stays unknown.

  • Search crawlers: permission to retrieve pagesOAI-SearchBot (ChatGPT search), Claude-SearchBot (Claude search) and PerplexityBot retrieve web content for search. A robots.txt restriction can limit retrieval; allowing a crawler does not guarantee indexing, citations, or customers.
  • Content-use policies: check each vendor’s scopeGPTBot and ClaudeBot are separate from their vendors’ search crawlers. Google-Extended is a product token for Gemini training and grounding, not a separate HTTP crawler. It does not control Google Search inclusion or ranking. Review Google’s documented scope before choosing that policy.
  • llms.txt: optional, and we say soWe check whether you publish one. Google says plainly you do not need one to appear in AI search. It is a cheap experiment some non-Google tools read, not a requirement, and we will not scare you about it.
  • The firewall behind the welcome matA scanner probe uses a documented bot user-agent from our address. An HTTP refusal is evidence about that request. It does not prove that the vendor’s verified crawler is blocked by your firewall.

If a bot you want is being turned away, the repair is in fixing AI crawler access and fixing robots.txt rules that block AI crawlers. For why any of this decides whether you appear in AI answers, start with what AEO is.

If you want to experiment with an optional llms.txt file, the llms.txt generator helps draft one. Prioritize actual crawl or indexing problems before optional files.

Crawler access is one part of search readiness. A GEO audit also reviews content structure and business identity. These checks describe the site; they do not establish whether an AI system will cite it.

crawler questions, answered straight

The honest version.

Which AI crawlers should I allow in robots.txt?

Allow the search-class crawlers if you want customers to find you through AI: OAI-SearchBot (ChatGPT search), Claude-SearchBot (Claude search) and PerplexityBot (Perplexity). Those are the ones that decide whether an engine can read and cite your pages in its answers. Training and content-use controls have different scopes. GPTBot and ClaudeBot are distinct from their vendors’ search crawlers. Google-Extended also controls Gemini grounding, while leaving Google Search inclusion and ranking unaffected.

Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot is OpenAI's training crawler; blocking it opts your content out of model training and nothing else. ChatGPT's search results come from a different crawler, OAI-SearchBot, which OpenAI documents as an independent control. The same split exists at Anthropic (ClaudeBot for training, Claude-SearchBot for search). Many sites copy a blanket block-everything robots.txt from a blog post and accidentally block the search crawlers too, which is the mistake this checker catches.

Do I need an llms.txt file to show up in AI answers?

No. Google's AI-search guidance says outright that you do not need new machine-readable files or AI text files to appear in generative AI search. An llms.txt is an optional, unofficial experiment: a short index of your key pages that a few non-Google AI tools read. It costs almost nothing to publish, so we report whether you have one, but no engine vendor requires it and skipping it does not hold you back with Google.

no email wall

Access is the first gate. Then 187 registered detectors.

Getting the crawlers through the door only matters if what they find is worth citing. The free whole-site scan evaluates applicable checks from the 187-detector catalog across SEO, answer eligibility, and AI readiness: category scores and selected findings, with missing measurements labeled, no email wall. Policy templates if you need to decide training vs search: AI bot policy. We make money on the keep-forever report and ongoing watch, not your inbox. Pricing. For AI-readiness beyond crawler access, answer signals, entity clarity, and citability, run a full AEO readiness audit for the complete picture.

Free category scores and selected findings. Full findings and export files: $10 once.