Statistics

SEO and AI Visibility Statistics

These statistics are drawn from our own scan data: AI crawler access rates across the Tranco top 1,000, AI citability checks across 105 scanned business sites, and per-bot policy data for 60 small-business sites. Every number is dated and sourced. None are invented or drawn from unverified external reports.

How many business sites open with a passage an AI engine can quote?

Only 3 of 105 scanned business sites (2.9%) opened with a passage that an AI engine could lift and quote verbatim. The check evaluated whether the first substantial paragraph after the main heading is self-contained, specific, and attributable: the qualities that make a passage usable by ChatGPT, Perplexity, or Google's AI features. Source: AuditLamp scan corpus, n=105, 2026-08-11.

Of the 99 sites that failed the answer-block check, 84 had no quotable passage on the homepage itself. The most common pattern was a slogan or hero line built to look good over a background image, with the first concrete fact about the business appearing three or more scrolls down. Source: AuditLamp scan corpus, n=105, 2026-08-11.

None of the sites in our corpus passed the answer-block check across a large multi-page scan. All three passing sites were single-page scans. Passing one page is easier than passing forty. Source: AuditLamp scan corpus, n=105, 2026-08-11.

What share of small business sites have decided what AI can see?

Zero of 60 small business sites blocked a search-class AI crawler in robots.txt on purpose. Not one owner had written a rule that would make them invisible to AI search answers. Source: AuditLamp scan corpus, n=60, 2026-07-15.

57 of 60 small business sites (95%) have no AI crawler rules in robots.txt at all. Their sites are open by default, not by any decision anyone made. Source: AuditLamp scan corpus, n=60, 2026-07-15.

3 of 60 small business sites (5%) opted out of AI training crawlers only. All three carried an identical Cloudflare-managed six-bot block, a deliberate choice that leaves AI search visibility untouched. Source: AuditLamp scan corpus, n=60, 2026-07-15.

How often does a site block AI at the edge without knowing it?

1 in 8 scanned business sites (13 of 102, 12.7%) returned HTTP 403 to an AI crawler even though their robots.txt explicitly allowed that crawler. The robots.txt said "welcome"; the server said "forbidden." The two layers disagreed, and in every case, the stricter one was the silent edge block the owner had never seen. Source: AuditLamp scan corpus, n=105 (check ran on 102), 2026-08-12.

Only 2 of 102 scanned sites blocked an AI crawler in robots.txt on purpose. The remaining 11 edge blocks were invisible to the owners -- a firewall setting, almost certainly a CDN default, that nobody had reviewed. Source: AuditLamp scan corpus, n=105 (check ran on 102), 2026-08-12.

11 of 60 small business sites (18%) refused a documented AI-bot user-agent at the network firewall with a 403 or 429, while a normal browser sailed through the same page without issue. Source: AuditLamp scan corpus, n=60, 2026-07-15.

7 of 60 small business sites (12%) turned away a search-class AI crawler (the kind that can cite them in answers people actually read) with no sign in robots.txt that anyone had chosen that outcome. Source: AuditLamp scan corpus, n=60, 2026-07-15.

How often are AI training crawlers blocked compared to retrieval crawlers?

OpenAI's training crawler, GPTBot, is blocked by 16.7% of the 653 reachable Tranco top-1,000 sites. OpenAI's search crawler, OAI-SearchBot -- the one that fetches pages live to answer a user question and may link them to the source -- is blocked by only 9.2%. The same company's two crawlers are refused at nearly a 2-to-1 ratio. Source: AuditLamp study, Tranco n=653, 2026-07-05.

Anthropic's training crawler, ClaudeBot, is blocked by 16.1% of reachable top-1,000 sites. Claude-User, the retrieval crawler that fetches pages to answer user queries, is blocked by 10.3%. The pattern holds across both major AI labs: train me less, cite me more. Source: AuditLamp study, Tranco n=653, 2026-07-05.

5.8% of reachable top-1,000 sites (38 sites) block at least one AI training crawler while explicitly allowing every AI retrieval crawler tested. This is a machine-readable "cite me, don't train me" policy: publishers willing to be quoted in an answer but unwilling to be free training data. Source: AuditLamp study, Tranco n=653, 2026-07-05.

49 of the reachable top-1,000 sites block GPTBot but explicitly allow OAI-SearchBot. Named sites include LinkedIn, Yahoo, Medium, Forbes, and eBay. These sites have made a deliberate distinction between OpenAI's training use and its retrieval use. Source: AuditLamp study, Tranco n=653, 2026-07-05.

Which AI crawler gets blocked most often at the top of the web?

The most-blocked AI-related crawler among the Tranco top-1,000 is not OpenAI or Anthropic. It is Common Crawl's CCBot, blocked by 18.7% of reachable sites. Common Crawl is an open dataset that feeds many downstream language models at once, so blocking it is the highest-leverage move for publishers who want to limit training use across multiple providers in a single robots.txt rule. Source: AuditLamp study, Tranco n=653, 2026-07-05.

ByteDance's Bytespider is blocked by 17.6% of reachable top-1,000 sites -- second only to CCBot. PerplexityBot, the indexing crawler for Perplexity's AI search product, is blocked by 14.1%. Perplexity-User, the live-fetch crawler that retrieves pages during a user query, is blocked by 11.2%. Source: AuditLamp study, Tranco n=653, 2026-07-05.

Do most sites block any AI crawler at all?

77.9% of the 653 reachable Tranco top-1,000 sites block no AI crawler at all. The AI-blocking conversation is loud, but even at the tech-forward head of the web, the majority of the internet is wide open to AI crawlers, on purpose or by default. Source: AuditLamp study, Tranco n=653, 2026-07-05.

Only 2.0% of reachable top-1,000 sites block Googlebot. Only 2.3% block Bingbot. A decade of SEO dependency on classic search has made these crawlers nearly untouchable, a pattern that has not changed despite the rise of AI search. Source: AuditLamp study, Tranco n=653, 2026-07-05.

How many major sites publish an llms.txt file?

11.9% of the 653 reachable Tranco top-1,000 sites publish an llms.txt file. Adopters include Cloudflare, GitHub, Azure, Shopify, WordPress, Adobe, Samsung, Dropbox, and PayPal, all at the tech-forward head of the web. In the long tail of ordinary business sites, adoption is near zero. Source: AuditLamp study, Tranco n=653, 2026-07-05.

What does research say about content formatting and AI citation lift?

A study measuring what actually lifts a page's visibility in generative-engine answers (GEO study, arXiv 2311.09735, later published at KDD 2024) found that adding attributed quotations produced approximately 41% position-adjusted visibility lift. The signals that worked were substance signals: statistics, quotations, and cited sources. Keywords and formatting tricks did not produce the same result. Source: GEO study (arXiv 2311.09735, KDD 2024), as cited in AuditLamp scan note, 2026-08-11.

How is this data sourced?

All statistics on this page come from one of two AuditLamp data sources. The Tranco top-1,000 study (measured 2026-07-05) fetched robots.txt and llms.txt from the 1,000 most-visited websites using the Tranco list JZ2VY (dated 2026-07-01), evaluated against RFC 9309 rules, and reported against the 653 sites that answered a plain HTTPS request. The full dataset is downloadable at auditlamp.com/study/ai-crawler-blocking. The scan-corpus figures come from real business sites that ran our audit engine; sample sizes and dates are noted per statistic. Neither source is a random sample of the whole web; rates in the long tail may differ.

What should a business owner do with these numbers?

Three things are worth acting on today, all free. First, check whether your site blocks an AI search crawler at the edge with a setting you never chose -- our free scan reads both layers and names the conflict in plain language. Second, read your homepage the way an AI engine does: does the first paragraph after your main heading contain a specific, attributable fact? If not, rewriting that one paragraph is the cheapest fix on this list. Third, if the training-versus-retrieval distinction matters to you, make it a written decision in robots.txt rather than an inherited CDN default. Run the free scan at auditlamp.com/app.

See where your site stands on these checks.

Paste your link. We read it the way Google and the AI engines do and print the failures in fix order. The preview is free.