SEO and AI Visibility Statistics
These statistics are drawn from our own scan data: AI crawler access rates across the Tranco top 1,000, AI citability checks across 105 scanned business sites, and per-bot policy data for 60 small-business sites. Every number is dated and sourced. None are invented or drawn from unverified external reports.
How many business sites open with a passage an AI engine can quote?
Only 3 of 105 scanned business sites (2.9%) opened with a passage that an AI engine could lift and quote verbatim. The check evaluated whether the first substantial paragraph after the main heading is self-contained, specific, and attributable: the qualities that make a passage usable by ChatGPT, Perplexity, or Google's AI features. Source: AuditLamp scan corpus, n=105, 2026-08-11.
Of the 99 sites that failed the answer-block check, 84 had no quotable passage on the homepage itself. The most common pattern was a slogan or hero line built to look good over a background image, with the first concrete fact about the business appearing three or more scrolls down. Source: AuditLamp scan corpus, n=105, 2026-08-11.
None of the sites in our corpus passed the answer-block check across a large multi-page scan. All three passing sites were single-page scans. Passing one page is easier than passing forty. Source: AuditLamp scan corpus, n=105, 2026-08-11.
What share of small business sites have decided what AI can see?
Zero of 60 small business sites blocked a search-class AI crawler in robots.txt on purpose. Not one owner had written a rule that would make them invisible to AI search answers. Source: AuditLamp scan corpus, n=60, 2026-07-15.
57 of 60 small business sites (95%) have no AI crawler rules in robots.txt at all. Their sites are open by default, not by any decision anyone made. Source: AuditLamp scan corpus, n=60, 2026-07-15.
3 of 60 small business sites (5%) opted out of AI training crawlers only. All three carried an identical Cloudflare-managed six-bot block, a deliberate choice that leaves AI search visibility untouched. Source: AuditLamp scan corpus, n=60, 2026-07-15.
How often does a site block AI at the edge without knowing it?
1 in 8 scanned business sites (13 of 102, 12.7%) returned HTTP 403 to an AI crawler even though their robots.txt explicitly allowed that crawler. The robots.txt said "welcome"; the server said "forbidden." The two layers disagreed, and in every case, the stricter one was the silent edge block the owner had never seen. Source: AuditLamp scan corpus, n=105 (check ran on 102), 2026-08-12.
Only 2 of 102 scanned sites blocked an AI crawler in robots.txt on purpose. The remaining 11 edge blocks were invisible to the owners -- a firewall setting, almost certainly a CDN default, that nobody had reviewed. Source: AuditLamp scan corpus, n=105 (check ran on 102), 2026-08-12.
11 of 60 small business sites (18%) refused a documented AI-bot user-agent at the network firewall with a 403 or 429, while a normal browser sailed through the same page without issue. Source: AuditLamp scan corpus, n=60, 2026-07-15.
7 of 60 small business sites (12%) turned away a search-class AI crawler (the kind that can cite them in answers people actually read) with no sign in robots.txt that anyone had chosen that outcome. Source: AuditLamp scan corpus, n=60, 2026-07-15.
How often are AI training crawlers blocked compared to retrieval crawlers?
OpenAI's training crawler, GPTBot, is blocked by 16.7% of the 653 reachable Tranco top-1,000 sites. OpenAI's search crawler, OAI-SearchBot -- the one that fetches pages live to answer a user question and may link them to the source -- is blocked by only 9.2%. The same company's two crawlers are refused at nearly a 2-to-1 ratio. Source: AuditLamp study, Tranco n=653, 2026-07-05.
Anthropic's training crawler, ClaudeBot, is blocked by 16.1% of reachable top-1,000 sites. Claude-User, the retrieval crawler that fetches pages to answer user queries, is blocked by 10.3%. The pattern holds across both major AI labs: train me less, cite me more. Source: AuditLamp study, Tranco n=653, 2026-07-05.
5.8% of reachable top-1,000 sites (38 sites) block at least one AI training crawler while explicitly allowing every AI retrieval crawler tested. This is a machine-readable "cite me, don't train me" policy: publishers willing to be quoted in an answer but unwilling to be free training data. Source: AuditLamp study, Tranco n=653, 2026-07-05.
49 of the reachable top-1,000 sites block GPTBot but explicitly allow OAI-SearchBot. Named sites include LinkedIn, Yahoo, Medium, Forbes, and eBay. These sites have made a deliberate distinction between OpenAI's training use and its retrieval use. Source: AuditLamp study, Tranco n=653, 2026-07-05.
Which AI crawler gets blocked most often at the top of the web?
The most-blocked AI-related crawler among the Tranco top-1,000 is not OpenAI or Anthropic. It is Common Crawl's CCBot, blocked by 18.7% of reachable sites. Common Crawl is an open dataset that feeds many downstream language models at once, so blocking it is the highest-leverage move for publishers who want to limit training use across multiple providers in a single robots.txt rule. Source: AuditLamp study, Tranco n=653, 2026-07-05.
ByteDance's Bytespider is blocked by 17.6% of reachable top-1,000 sites -- second only to CCBot. PerplexityBot, the indexing crawler for Perplexity's AI search product, is blocked by 14.1%. Perplexity-User, the live-fetch crawler that retrieves pages during a user query, is blocked by 11.2%. Source: AuditLamp study, Tranco n=653, 2026-07-05.
Do most sites block any AI crawler at all?
77.9% of the 653 reachable Tranco top-1,000 sites block no AI crawler at all. The AI-blocking conversation is loud, but even at the tech-forward head of the web, the majority of the internet is wide open to AI crawlers, on purpose or by default. Source: AuditLamp study, Tranco n=653, 2026-07-05.
Only 2.0% of reachable top-1,000 sites block Googlebot. Only 2.3% block Bingbot. A decade of SEO dependency on classic search has made these crawlers nearly untouchable, a pattern that has not changed despite the rise of AI search. Source: AuditLamp study, Tranco n=653, 2026-07-05.
How many major sites publish an llms.txt file?
11.9% of the 653 reachable Tranco top-1,000 sites publish an llms.txt file. Adopters include Cloudflare, GitHub, Azure, Shopify, WordPress, Adobe, Samsung, Dropbox, and PayPal, all at the tech-forward head of the web. In the long tail of ordinary business sites, adoption is near zero. Source: AuditLamp study, Tranco n=653, 2026-07-05.
What does research say about content formatting and AI citation lift?
A study measuring what actually lifts a page's visibility in generative-engine answers (GEO study, arXiv 2311.09735, later published at KDD 2024) found that adding attributed quotations produced approximately 41% position-adjusted visibility lift. The signals that worked were substance signals: statistics, quotations, and cited sources. Keywords and formatting tricks did not produce the same result. Source: GEO study (arXiv 2311.09735, KDD 2024), as cited in AuditLamp scan note, 2026-08-11.
How is this data sourced?
All statistics on this page come from one of two AuditLamp data sources. The Tranco top-1,000 study (measured 2026-07-05) fetched robots.txt and llms.txt from the 1,000 most-visited websites using the Tranco list JZ2VY (dated 2026-07-01), evaluated against RFC 9309 rules, and reported against the 653 sites that answered a plain HTTPS request. The full dataset is downloadable at auditlamp.com/study/ai-crawler-blocking. The scan-corpus figures come from real business sites that ran our audit engine; sample sizes and dates are noted per statistic. Neither source is a random sample of the whole web; rates in the long tail may differ.
What should a business owner do with these numbers?
Three things are worth acting on today, all free. First, check whether your site blocks an AI search crawler at the edge with a setting you never chose -- our free scan reads both layers and names the conflict in plain language. Second, read your homepage the way an AI engine does: does the first paragraph after your main heading contain a specific, attributable fact? If not, rewriting that one paragraph is the cheapest fix on this list. Third, if the training-versus-retrieval distinction matters to you, make it a written decision in robots.txt rather than an inherited CDN default. Run the free scan at auditlamp.com/app.