Cloudflare's Search, Training and Agent bot settings: what they change in your robots.txt
If your site is on Cloudflare, your robots.txt may no longer be only the file you wrote. Here is what the 2026 settings do, which one keeps you in search, and how to check what crawlers are actually being asked.
The three controls
Since July 1, 2026, Cloudflare has let you set a policy for three kinds of bot traffic, per domain (Cloudflare, August 21; definitions from Cloudflare, September 15):
- Search: crawling to build a search index.
- Training: crawling to train or fine-tune a model.
- Agent: user-directed agents visiting a page for a person, such as chat fetch bots and browser-use agents.
Search and Agent can be set to Allow, Block on pages with ads, or Block. Training adds a fourth option, Disallow AI Training (Cloudflare).
What changed on September 15, 2026
Some crawlers do search and training under one user agent. Cloudflare calls them mixed-use crawlers and names Applebot, Bingbot and Googlebot. From September 15, 2026 (Cloudflare):
- Block and Block on pages with ads for Training now also apply to those mixed-use crawlers, so either setting affects search as well as training.
- Disallow AI Training publishes a no-training preference in robots.txt and keeps the crawlers Cloudflare labels Accountable (Apple, Google and Microsoft) allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta and OpenAI.
- The old Block AI Bots switch is being deprecated in favour of the three controls, and Managed Robots.txt is being replaced by Bot Preference Sync.
- There is no Disallow option for Agents. Cloudflare says the web has no well-established directive for that yet.
Existing settings were migrated. A domain that used the legacy Block AI Bots switch (Block, or Block on pages with ads) moved to Search Allow, Training Disallow AI Training, and Agent Block on pages with ads. A Training setting of Block or Block on pages with ads that was configured before September 15 moved to Disallow AI Training (Cloudflare). So if you want Googlebot, Bingbot and Applebot blocked entirely, you now have to choose Block yourself, and that removes you from their search too.
Which setting does what
| Your goal | Training setting | Effect, per Cloudflare |
|---|---|---|
| Stay in search, keep content out of AI training | Disallow AI Training | A no-training preference is written to robots.txt. Accountable mixed-use crawlers keep crawling for search; other training crawlers are blocked. |
| Block training crawlers everywhere, search included | Block | All training crawlers are blocked, including Applebot, Bingbot and Googlebot, so search is affected. |
| Block training crawlers only where you show ads | Block on pages with ads | Crawlers, mixed-use ones included, are blocked only on pages detected as serving an ad. Cloudflare offers no Disallow AI Training for ad pages, because that page list can't be written into robots.txt. |
| No restriction | Allow | All crawlers allowed unless another setting or a WAF rule blocks them. |
For new domains, Cloudflare offers presets. If you say the site is monetised with ads, the preset is Search Allow, Training Disallow AI Training and Agent Block on pages with ads. Otherwise everything is set to Allow. Bot Preference Sync is enabled in both presets (Cloudflare).
How the no-training preference reaches each search crawler
- Google: through a Disallow for the
Google-Extendedtoken (Cloudflare). Google says Google-Extended does not affect a site's inclusion in Google Search and is not used as a ranking signal (Google). - Apple: through a Disallow for
Applebot-Extended(Cloudflare). Apple says that even if you disallow Applebot-Extended, your content remains discoverable through Spotlight, Siri and Safari (Apple). - Bing: not through robots.txt yet. Cloudflare reports that Microsoft is building support for a robots.txt no-training preference, targeted for early 2027, and that until then Disallow AI Training does not pass the preference to Bing. Bing's current mechanism is the
NOARCHIVEmeta tag (Cloudflare).
Bot Preference Sync writes to your robots.txt
Bot Preference Sync turns your dashboard choices into robots.txt rules. It is available on every plan, can be turned on or off at any time, and is on by default for new customers. If you already have a robots.txt, its lines are prepended to your file between # BEGIN Cloudflare Bot Preference Sync and # END Cloudflare Bot Preference Sync, and your own rules stay below. Cloudflare updates the list of bots over time. It does not read your individual custom rules, so if you make exceptions there, you can turn the sync off and write the file yourself (Cloudflare).
Cloudflare's example of the added block, shortened by Cloudflare:
# BEGIN Cloudflare Bot Preference Sync User-agent: TrainingBot1 User-agent: TrainingBot2 User-agent: TrainingBot3 User-agent: MixedUseBot-Extended Disallow: / ... # END Cloudflare Bot Preference Sync
Watch for overlaps with your own rules
Because the synced block sits above your own rules in the same file, both are read together. Under RFC 9309, if more than one group names a crawler, the groups are combined into one, and the longest matching rule wins (RFC 9309, 2.2.1 and 2.2.2). For example:
# added by the sync User-agent: GPTBot Disallow: / # your own, older rules further down User-agent: GPTBot Allow: /blog/
GPTBot now gets Disallow: / and Allow: /blog/ as one group, and /blog/ is the longer match, so the file still allows it under /blog/. The file is not the whole story, though: with Disallow AI Training or Block, Cloudflare also blocks training-only crawlers such as OpenAI's at its edge (Cloudflare), whatever robots.txt says. Clean up overlapping groups so the file says what you mean.
Also remember that a crawler named in any group ignores the User-agent: * group. If the synced block names a crawler, your * rules no longer apply to it.
Check what your file says now
- Open
https://yourdomain/robots.txtand look for the# BEGIN Cloudflare Bot Preference Syncblock. - Run the RobotsGate site check on your domain. For every AI crawler in the registry, it shows allowed or blocked for a path, the group that matched, and the line that decided it, with groups combined as RFC 9309 requires.
- Check a few paths that matter to you, such as
/and/blog/, with thepathoption.
The check reads robots.txt only. It can't see Cloudflare's edge blocks, WAF rules or page-level tags such as NOARCHIVE. See How to check which AI bots can crawl your site for what a robots.txt check can and can't tell you.