Infrastructure

Cloudflare Bot Preference Sync writes your robots.txt, and its three settings cannot express a per-crawler policy

September 18, 2026 at 4:25 PM PT

Search Engine Journal illustration for Cloudflare Bot Preference Sync

Image: Search Engine Journal

Why it mattersA team whose robots.txt gets managed by Cloudflare needs to check what the file now says. If your policy is per company rather than per category, the sync will overwrite it with something you did not choose.

Cloudflare Bot Preference Sync writes a website's robots.txt from three settings in the Cloudflare dashboard. Chris Green, writing in Search Engine Journal on 18 September, argues that the three categories cannot express what many websites actually do, and that Cloudflare will turn the sync on by default for new customers.

Cloudflare announced Bot Preference Sync on 21 August. The product prepends generated robots.txt entries between # BEGIN Cloudflare Bot Preference Sync and # END Cloudflare Bot Preference Sync markers, leaving your existing file underneath. Cloudflare says it will run from the free tier up. On 15 September, Cloudflare said the default for new domains will be to block Training and Agent on pages that show ads, and leave Search allowed. Onboarding asks whether you monetize the site with ads, and the answer becomes a published position on AI training.

The three settings and what they cannot say

Bot Preference Sync generates the file from Search, Agent, and Training, each with the same options: block on all pages, block on pages with ads, or allow. Training also has a Disallow option that writes a no-training line into the file. Cloudflare's tracked bot list decides which crawlers fall into which category. Green writes that his own policy allows OpenAI's GPTBot, Anthropic's crawler, and PerplexityBot, and blocks Bytespider and meta-externalagent, because the first three send readers back and the other two do not. Every company on both sides trains models, so no single setting for Training expresses his rule. Cloudflare's documented answer is to turn the sync off and maintain the file by hand.

The four disclosure conditions attached to Disallow Training

Cloudflare has published four conditions an AI crawler must meet to avoid being treated as opaque when a site sets Training to Disallow: it must respect a no-training preference in robots.txt, provide a way to opt out of AI summaries, give URL-level visibility into pages used for training and search metrics, and show publicly that disallowing training does not hurt search results. Miss any of them and Cloudflare blocks the crawler on every site that set Training to Disallow.

Green points out that Google meets condition four through its Google-Extended documentation, updated 14 July 2026, which states the control does not affect Google Search ranking. Google has no answer for condition two: its AI features documentation names nosnippet, data-nosnippet, max-snippet, and noindex as the only controls, and all of them affect ordinary search snippets. Microsoft answered condition two in September 2023 with NOARCHIVE, which keeps content out of Bing Chat answers while keeping it in the index.

What robots.txt actually enforces

Green measured the crawler traffic on his own site and reported that the largest so-called AI crawler in his logs was asking for /.env and SSH keys under a nonprofit research archive's name. A line in a text file does not stop a crawler that never intended to obey it. Bot Preference Sync is useful for the crawlers that do choose to obey, which Green calls the honest half of the internet. Enforcement stays at the edge; the file records intent.

Cloudflare sits in front of a large share of the web, and now writes what many websites will publish to AI crawlers. The conditions attached to Training are the vendor's, the enforcement is the vendor's network, and the sites doing the blocking mostly clicked one toggle. Ten minutes reading your current robots.txt against your Cloudflare AI bot policies tells you whether the sync would change your published position, and whether the three categories can carry your actual rule.

Source

Reported by: Search Engine Journal

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

Google, Cloudflare, and Microsoft are running three different payment models for AI use of publisher content, and each measures a different event

Matt G. Southern's 18 September piece in Search Engine Journal sets Google's AI contribution pilot, Cloudflare's Pay Per Crawl and Pay Per Use, and Microsoft's Publisher Content Marketplace against each other. Each pays for a different event, gives site owners a different level of control, and returns different data.

Source: PressGo-to-market

Suganthan Mohanadasan set his site to charge AI agents one cent per page, and Claude Code paid it from a wallet during a task

SEO consultant Suganthan Mohanadasan set his site to return HTTP 402 to AI agents and release a page for one cent in testnet USDC, and reports five settled crawls on 15 September, including one paid by Claude Code from a wallet during a task.

Source: PressGo-to-market

Cloudflare adds a Disallow AI Training setting that keeps Google, Apple and Bing crawling for search

Cloudflare has split its bot controls so a site can opt out of AI training while keeping Googlebot, Applebot and Bingbot crawling for search, reversing an earlier plan that would have blocked all three from Cloudflare-fronted sites that refused training.

Source: PressGo-to-market