For publishers · blogs · media sites

Can AI read your articles?

Most publishers have not made an active choice about AI crawler access: 95% of 60 small business sites in our corpus had no AI crawler rules in robots.txt at all. For a publication, that means your access policy is set by the platform default, not by a choice you made. They are open by default, not by any decision anyone made. Separately, 18% of those same sites refused a named AI crawler at the network level while normal browser traffic passed through fine. For a publisher, a network-level block means AI answer engines cannot read articles that were written to be cited. Source: AuditLamp scan corpus, n=60, 2026-07-15.

The free scan checks both layers: robots.txt policy and actual edge behavior for named AI bots. It also checks whether your pages carry quotable passages and a valid llms.txt.

Free whole-site scan · 164 graded checks · AI crawlers by name · no email on the diagnosis

Questions publishers ask

Do my robots.txt settings actually control what AI crawlers can access?

Often, no. In our scan corpus, 12.7% of 102 scanned sites returned HTTP 403 to a named AI crawler even though robots.txt explicitly allowed it. The CDN or edge firewall overrode the robots file. You can grant access in robots.txt and still be invisible to GPTBot at the network layer. For a publisher, that means articles you intended to be cited are unreachable to the AI systems making citation decisions. Source: AuditLamp scan corpus, n=105 (check ran on 102), 2026-08-12.

This is the most common AI visibility gap we see in publisher sites: a robots.txt that looks open and an edge configuration that silently closes the door. The AuditLamp scan probes each named crawler by user-agent separately from the robots.txt parse, so it can tell you whether the two layers agree. When they disagree, the stricter one wins and the owner usually does not know.

Is llms.txt necessary for a blog or media site?

Not strictly required. AI crawlers work from robots.txt and direct page fetches and do not need llms.txt to access your content. But llms.txt lets you point AI systems at your best pages and exclude low-value content. It is a guidance file, not a gate. Absence does not block access; presence shapes what gets indexed and cited first.

If your publication has hundreds of pages, llms.txt is a way to tell an AI what is worth reading: your research, your primary sources, your high-signal content. Without it, the crawler makes that judgment itself. The scan validates your existing llms.txt format and flags broken links inside it. The llms.txt generator can draft one from your sitemap if you do not have one yet.

How do I get my content cited by ChatGPT or Perplexity?

Access alone does not produce citations. A crawler must fetch the page, find a self-contained quotable passage, and confirm the specific AI user-agent is not blocked. Only 2.9% of 105 scanned business sites opened with a passage an AI engine could lift and quote verbatim. For a publication, that means most pages AI crawlers can reach do not produce citations because the content is not structured as a quotable passage. Source: AuditLamp scan corpus, n=105, 2026-08-11.

The quotability check is separate from the access check. A page can be fully accessible to GPTBot and still fail because the first substantial content block after the heading is a slogan, a photo caption, or a navigation element. AI answer engines need a passage that is specific, self-contained, and attributable: the kind of thing you can lift out of context and it still makes sense. The scan flags which pages pass this check and which do not.

Why is an AI overview not available for my content on Google?

AI overviews require that your page be crawled, indexed, and contain a passage answering a specific query in one extractable block. Common blockers: AI crawlers refused at the edge despite robots.txt permission, content that requires JavaScript to render, and pages where the first concrete information appears too far below the heading for the crawler to prioritize it.

The JavaScript rendering problem is particularly common on CMS-based publications. If your article body loads via client-side rendering and the crawler receives an empty shell, Google's systems see no content to feature. The scan checks whether pages return readable content in raw HTML or require script execution to display anything, and flags the specific pages where this is a problem.

What makes a page quotable to an AI answer engine?

A quotable page has a self-contained passage in the first substantial section after the heading, specific enough to stand alone without surrounding context, accessible to AI crawlers without JavaScript rendering. Only 2.9% of 105 scanned business sites passed this check on their homepage. For a publisher, a homepage that fails this check signals to AI engines that your site is not a citable source, regardless of what is deeper in the archive. Source: AuditLamp scan corpus, n=105, 2026-08-11.

The difference between a passing and failing page is usually not content length. It is placement and specificity. A page that opens with "Welcome to our publication, where we cover everything" fails. A page that opens with "GPTBot is blocked by 16.7% of Tranco top-1,000 sites, according to AuditLamp's July 2026 crawler study" passes. The scan grades your actual first passage and tells you what needs to change.

Related tools

What we do not do, and what fills the gap.

We are not a live citation tracker for ChatGPT or Perplexity answers: that is a different, expensive product class. We measure readiness and access on your site, the half you control. Method notes are public in the learn library and crawler study. Free tools: AI crawler access checker, llms.txt generator.

Check your publication's AI crawler access.

AI startups · saas · designers · local · study