Cloudflare lets sites refuse AI training without giving up search, and names Apple, Google and Microsoft as Accountable

Image: Cloudflare
Why it mattersThe all-or-nothing block that stopped a lot of sites opting out of AI training is over on Cloudflare, so a site that wants Google Search but not Gemini training has a supported way to say so.
Cloudflare published a new setting today called Disallow AI Training, which publishes robots.txt directives that refuse AI model training while leaving search crawling in place. Alongside the setting, Cloudflare named Apple, Google and Microsoft as Accountable crawler operators, meaning their AI bots honour the split.
What the setting does
The setting publishes per-bot directives in robots.txt, aimed at the identifiers those companies already use for training crawlers: Applebot-Extended, Google-Extended, and the equivalents from Microsoft. A site owner who turns it on continues to be crawled by Googlebot, Bingbot and Applebot for search, and stops being crawled by Google-Extended and Applebot-Extended for model training. Cloudflare says the three behaviours, search, training and agent access, are now three separate controls on the same domain rather than one block-everything switch.
The Accountable designation
To qualify as an Accountable operator, Cloudflare says a crawler owner must meet four requirements: honour training opt-outs, honour opt-outs for AI summaries, provide URL-level visibility into what was used for training, and confirm that refusing training will not hurt search ranking on the operator's own surface. Apple, Google and Microsoft qualified. The Cloudflare post notes that each combines capabilities available today with time-bound commitments for parts still in development, so the designation is partly forward-looking.
The number that made this necessary
Cloudflare says 17% of sites now enable some mechanism to block training, while fewer than 1% block search bots. That is the case for granular control in one sentence: a lot of sites want to refuse training and almost none want to refuse search, and until now the practical choice on many surfaces was refuse both or refuse neither. New domains on Cloudflare get preset defaults keyed to how the site makes money, and ad-supported sites default to Disallow AI Training with search left open.
Why this changes the picture
Refusing training used to cost search visibility on Google specifically. Google-Extended was framed as decoupling training from search, but a lot of publishers still went the whole-hog route because they did not trust the split to hold and did not want to test it against ranking. Getting Google itself onto the Accountable list, on Cloudflare's terms and with Cloudflare's audit of the four requirements, moves the split from a promise on a docs page to a claim a large intermediary is willing to name and enforce with directives.
For anyone whose GEO strategy has been to be quotable by ChatGPT, Perplexity and Google AI Overviews while refusing to be fine-tuned on, this is the first day the two positions do not fight each other on Cloudflare. Set the training opt-out, leave search on, and the retrieval crawlers still index the site so it can be cited when a live query lands. The Cloudflare post says the setting is available now to all sites on the platform.
Source
Source: Cloudflare, 15 September 2026.
Source: Cloudflare
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually shipping with them. Short, and only when there is something worth reading.


