Go-to-market

Cloudflare adds a Disallow AI Training setting that keeps Google, Apple and Bing crawling for search

September 16, 2026 at 12:20 PM PT

Cloudflare diagram of the new Training, Search and Agent bot controls with Accountable mixed-use crawlers routed around the Disallow AI Training block

Image: Cloudflare

Why it mattersAny site behind Cloudflare that was heading into a full search blackout on 15 September now gets a middle setting that keeps the search index alive, and the choice is a dashboard toggle rather than a robots.txt rewrite.

Cloudflare has launched a new Disallow AI Training setting on the Training control it ships to every zone, and it says sites that already block AI training will be migrated onto it. The setting publishes a no-training preference in robots.txt while still letting Googlebot, Applebot and Bingbot crawl the site for search. It went live on 15 September.

The change reverses a plan Cloudflare described in July, under which sites that blocked AI training on 15 September would also lose those three crawlers, because each is used for both search and training. Cloudflare says most customers will see no change: existing Training selections of Block or Block on pages with ads will move to Disallow AI Training, and Block AI Bots and the Managed Robots.txt feature are being retired.

What the three controls now do

Cloudflare classifies bots by behaviour and offers three independent controls: Search, Training and Agent. Training now has four settings. Allow lets every crawler in. Disallow AI Training publishes a Disallow for the training-only crawlers run by Amazon, Anthropic, Meta and OpenAI, and keeps mixed-use crawlers from labelled companies in for search. Block on pages with ads blocks every training crawler, including mixed-use ones, on pages that carry an ad. Block stops all training crawlers, including for search, which is the setting Cloudflare says now cuts a site off from Google, Apple and Bing entirely.

The Accountable label decides which crawlers stay in

Cloudflare has introduced a designation it calls Accountable and applied it to crawlers whose operators meet or commit to four requirements: a robots.txt opt-out for training, an opt-out for AI summaries, URL-level visibility into what was made available for training, and an assurance that opting out will not affect search ranking. Cloudflare says Apple, Google and Microsoft all qualify. Amazon, Anthropic, Meta and OpenAI qualify too, because they run separate search and training crawlers rather than a mixed-use one, so their training crawlers are blocked under Disallow AI Training without touching search.

How the opt-out reaches each engine

At Google, Disallow AI Training publishes a Disallow for Google-Extended, the robots.txt token Google offers for opting out of Gemini training. Google's own documentation says Google-Extended does not affect inclusion in Google Search or ranking. AI Overviews and AI Mode are controlled separately in Search Console.

At Apple, the setting publishes a Disallow for Applebot-Extended. Apple documents this as not affecting search ranking, and says keeping content out of Siri and Search AI answers takes the nosnippet meta tag on top of it.

Bing does not yet honour a robots.txt no-training preference, so Disallow AI Training sends nothing to Microsoft for now. Cloudflare says Microsoft's support for the directive is targeted for early 2027. In the meantime the training opt-out on Bing is the NOARCHIVE meta tag, which also removes links from Chat and Copilot.

Cloudflare says the next piece it is working on is a single AI-summaries control, so a site owner can set one preference for how much of its content ends up in an AI answer instead of tuning it per operator.

Source

Source: Cloudflare

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

SEJ column shows the AI crawl-to-refer ratio for Anthropic has been quoted at seven different values in 13 months, all from Cloudflare

Duane Forrester traces seven different versions of Anthropic's crawl-to-refer ratio, all attributed to Cloudflare and published inside 13 months, and shows the metric has four hidden denominators that never travel with the number.

Source: PressGo-to-market

Common Crawl read 584,107 llms.txt files, and the sites that wrote CCBot as blocked inside them do not block CCBot in robots.txt

Common Crawl analyzed 584,107 llms.txt files from its July 2026 crawl and reports that two thirds come from a plugin, 22.56% carry no links at all, and the sites naming CCBot as blocked in the file do not actually block it in robots.txt.

Source: PressGo-to-market

Apple says robots.txt rules for Applebot-Extended do not affect search ranking

Apple added one line to its Applebot documentation stating that robots.txt rules for Applebot-Extended, its AI training opt-out agent, are not considered in ranking for Apple's search results.

Source: PressGo-to-market