Common Crawl read 584,107 llms.txt files, and the sites that wrote CCBot as blocked inside them do not block CCBot in robots.txt

Image: Common Crawl
Why it mattersA team that wrote crawler rules into llms.txt has not blocked anything. The only file crawlers are obliged to read for that purpose is robots.txt, and that is where any real opt-out has to sit.
Common Crawl published a content analysis of 584,107 llms.txt files pulled from its July 2026 crawl, and the headline finding is that the file the AI search community has been treating as a robots.txt for language models does not do what people are writing into it. The post, by Common Crawl Senior Research Engineer Malte Ostendorff, went up on 31 August 2026 and reached wider circulation this week through a Search Engine Journal write-up.
What Common Crawl actually looked at
Common Crawl seeded /llms.txt and /llms-full.txt into its July 2026 crawl for a large sample of hosts, then filtered the 598,298 responses down to the 584,107 files that carried real content. Adoption in that sample was 11.72% for /llms.txt among hosts Common Crawl can fetch, well above the 2% the Web Almanac reported for 2025 across all sites. Common Crawl notes the two populations differ, so the numbers cannot be compared directly.
The shape travels, the substance does not
Common Crawl reports that 49.90% of the corpus follows the full specification: an H1, a summary blockquote, and section headings of link bullets. But 22.56% carry no links at all, in a file that is supposed to publish a curated list of URLs. Only 32.94% of the links that do exist carry the short note the spec requires. 68.27% of the corpus is templated, with Wix alone producing 41.34% of the files. All in One SEO writes 73,136 llms.txt files with a median of 138 links and never gets as far as a summary: the plugin treats the file as a sitemap. 2.54% of the corpus is a GoDaddy parked-domain boilerplate, a sales pitch aimed at a language model that still passes every structural test.
The finding a publisher should read twice
Common Crawl found 1,570 files (0.27%) that name a specific crawler and either allow or deny it. Of those, 32 files claim to block CCBot. On 17 August 2026, Common Crawl fetched robots.txt for all 32 of those sites and reports that none of them blocks CCBot in robots.txt. Five allow it explicitly, eleven allow it with a handful of excluded paths, and fifteen do not restrict it at all. Common Crawl quotes proform.com as the pattern: an llms.txt section headed "AI Crawler Access (robots.txt status as of June 2026)" that lists CCBot under Blocked, while the actual robots.txt admits CCBot under User-agent: * with Allow: /. Common Crawl restates its position: "llms.txt is not an access-control mechanism, and nothing obliges any crawler to read it."
Common Crawl also reports on prompt injection in the corpus. Ten files match its strictest test, and the team read all ten by hand: four are genuine injections, one is a bug-bounty researcher's deliberate payload, one asks the reader to become a catgirl, one is a joke about poisoning crawler logs, and three are false positives on the phrase "you are now" or on the literal ChatML tokens shown in a technical blog. Every real injection was placed there deliberately by someone technical enough to be making a point.
For a team working on being cited by answer engines, the practical takeaway is that any crawler rule written into llms.txt is decorative. Rules that a crawler will actually honor belong in robots.txt, and any opt-out written anywhere else has not been made. The llms.txt file is still useful as a curated map of pages worth reading, provided the links and their notes are real and the team maintaining it treats it as documentation.
Source
A Content Analysis of llms.txt Files from the July 2026 Crawl Archive, by Malte Ostendorff at Common Crawl, 31 August 2026. Reported by Search Engine Journal, 8 September 2026.
Source: Common Crawl
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually shipping with them. Short, and only when there is something worth reading.


