Note

llms.txt adoption: what 105 scanned sites actually have

The short answer: we pulled the latest scan for each of the 105 domains that have run our engine and looked at one check: is an llms.txt file published at the domain root? The check ran on 103 of the 105 sites, and 45 of them had one, about 44 percent. That is nearly four times the 11.9 percent we measured on the open web's most-visited sites, and the gap is not a mystery: these 105 domains chose to run an AI-visibility scan, so this corpus skews heavily toward owners already paying attention to AI search. Treat 44 percent as adoption among the AI-aware, not the internet.

What we found

| llms.txt result | Sites | |---|---| | File published at /llms.txt | 45 | | No file found (not graded as a defect) | 58 | | Check did not run | 2 | | Total | 105 |

One reading note on that table. Our engine does not treat a missing llms.txt as a failure, because no major engine requires the file. A site with no llms.txt gets an informational note, not a penalty, and that is deliberate. The 45 that published one get credit for a cheap, optional experiment, nothing more.

Why is this number so much higher than the open web?

Selection. In July 2026 we crawled the 1,000 most-visited websites on the Tranco list and found an llms.txt on 11.9 percent of the 653 reachable sites. The 105 domains in this study are different animals: businesses and site owners who deliberately ran a scan that reads their site the way AI engines do. People who do that have usually already read a post telling them to publish an llms.txt, and many followed the advice. When roughly 44 percent of an AI-curious group has the file and roughly 12 percent of the general top of the web does, the difference is the audience, not a trend line. Do not quote our 44 percent as web-wide adoption. It is not.

What do the published files actually look like?

This is where the data gets interesting. Across the 45 published files, the median size was 4,384 bytes, a page or two of curated links and descriptions, which matches the file's stated purpose. But the spread is wide: the smallest was 652 bytes and the largest was 253,038 bytes, about a quarter of a megabyte of text. Five of the 45 files exceeded 100 kilobytes.

A 250 KB llms.txt is almost certainly not a curated summary. It is a plugin or generator dumping every URL on the site into the file. The llms.txt proposal, such as it is (the format is an informal convention from llmstxt.org, with no official specification and no documented engine support), describes a short, structured orientation document for language models. If your file is a full sitemap in disguise, you have shipped the letter of the idea and skipped the point. For what the file can and cannot do, our guide on what an llms.txt file actually does for AI visibility is the honest version.

One more small number: newer scans also inspect the llms.txt body for a reference to the optional llms-full.txt companion file. Of the 34 files inspected, exactly one mentioned it.

Does Google read llms.txt?

No. Google's AI guidance (developers.google.com/search/docs/fundamentals/ai-optimization-guide) says verbatim that you do not need to create new machine-readable files, AI text files, markup, or Markdown to appear in Google Search, including its generative features. That sentence is why our engine scores a missing llms.txt as informational rather than a defect, and why we will not tell you the file is a ranking lever. For engines other than Google, there is no published evidence that the file changes citation frequency either way. It is a low-cost bet on the non-Google engines that do fetch it, and that is the whole claim.

Should you publish one?

If it takes you ten minutes, sure. The file has no proven downside, some non-Google AI tools read it, and it forces a useful exercise: deciding which of your pages actually matter and describing them in plain sentences. Keep it short and curated, closer to the 4 KB median than the 250 KB outlier. If you want a starting point, our llms.txt generator builds one from your site's pages. Just do not publish it instead of fixing the things engines demonstrably use: crawlable pages, real answers at the top of the page, and accurate structured data.

What this corpus is and is not

These 105 domains are self-selected: owners who ran our scan, skewed toward small firms, local services, and people already thinking about AI search. We took the latest scan per domain from our production scan records, read one check per site, and counted. This is not a random sample of the web and we make no claim about global adoption; for that, use our Tranco top-1,000 study. What this corpus does show is the ceiling of the advice cycle: even among owners actively working on AI visibility, fewer than half have the file, and among those who do, a meaningful slice have auto-generated dumps rather than the curated summary the format describes.

See where your site stands

Whether you have an llms.txt, what is in it, and whether the rest of your site gives AI engines anything worth citing are three different questions, and only the last one is the one that pays. Run a free Visibility Scan at auditlamp.com. It checks your llms.txt along with everything the engines actually read, and tells you which fixes matter in plain language.

Stop reading. Start with your own site.

Paste your link. We read it the way Google and the AI engines do and print the failures in fix order. The preview is free.