RobotsGate

Cloudflare's Search, Training and Agent bot settings: what they change in your robots.txt

If your site is on Cloudflare, your robots.txt may no longer be only the file you wrote. Here is what the 2026 settings do, which one keeps you in search, and how to check what crawlers are actually being asked.

This page summarises Cloudflare's own posts of August 21, 2026 and September 15, 2026, read on 2026-10-05. RobotsGate is not affiliated with Cloudflare. Cloudflare can change these settings at any time, so check its posts and dashboard for the latest. robots.txt is voluntary (RFC 9309).

The three controls

Since July 1, 2026, Cloudflare has let you set a policy for three kinds of bot traffic, per domain (Cloudflare, August 21; definitions from Cloudflare, September 15):

Search and Agent can be set to Allow, Block on pages with ads, or Block. Training adds a fourth option, Disallow AI Training (Cloudflare).

What changed on September 15, 2026

Some crawlers do search and training under one user agent. Cloudflare calls them mixed-use crawlers and names Applebot, Bingbot and Googlebot. From September 15, 2026 (Cloudflare):

Existing settings were migrated. A domain that used the legacy Block AI Bots switch (Block, or Block on pages with ads) moved to Search Allow, Training Disallow AI Training, and Agent Block on pages with ads. A Training setting of Block or Block on pages with ads that was configured before September 15 moved to Disallow AI Training (Cloudflare). So if you want Googlebot, Bingbot and Applebot blocked entirely, you now have to choose Block yourself, and that removes you from their search too.

Which setting does what

Your goalTraining settingEffect, per Cloudflare
Stay in search, keep content out of AI trainingDisallow AI TrainingA no-training preference is written to robots.txt. Accountable mixed-use crawlers keep crawling for search; other training crawlers are blocked.
Block training crawlers everywhere, search includedBlockAll training crawlers are blocked, including Applebot, Bingbot and Googlebot, so search is affected.
Block training crawlers only where you show adsBlock on pages with adsCrawlers, mixed-use ones included, are blocked only on pages detected as serving an ad. Cloudflare offers no Disallow AI Training for ad pages, because that page list can't be written into robots.txt.
No restrictionAllowAll crawlers allowed unless another setting or a WAF rule blocks them.

For new domains, Cloudflare offers presets. If you say the site is monetised with ads, the preset is Search Allow, Training Disallow AI Training and Agent Block on pages with ads. Otherwise everything is set to Allow. Bot Preference Sync is enabled in both presets (Cloudflare).

How the no-training preference reaches each search crawler

Bot Preference Sync writes to your robots.txt

Bot Preference Sync turns your dashboard choices into robots.txt rules. It is available on every plan, can be turned on or off at any time, and is on by default for new customers. If you already have a robots.txt, its lines are prepended to your file between # BEGIN Cloudflare Bot Preference Sync and # END Cloudflare Bot Preference Sync, and your own rules stay below. Cloudflare updates the list of bots over time. It does not read your individual custom rules, so if you make exceptions there, you can turn the sync off and write the file yourself (Cloudflare).

Cloudflare's example of the added block, shortened by Cloudflare:

# BEGIN Cloudflare Bot Preference Sync

User-agent: TrainingBot1
User-agent: TrainingBot2
User-agent: TrainingBot3
User-agent: MixedUseBot-Extended
Disallow: /

...

# END Cloudflare Bot Preference Sync

Watch for overlaps with your own rules

Because the synced block sits above your own rules in the same file, both are read together. Under RFC 9309, if more than one group names a crawler, the groups are combined into one, and the longest matching rule wins (RFC 9309, 2.2.1 and 2.2.2). For example:

# added by the sync
User-agent: GPTBot
Disallow: /

# your own, older rules further down
User-agent: GPTBot
Allow: /blog/

GPTBot now gets Disallow: / and Allow: /blog/ as one group, and /blog/ is the longer match, so the file still allows it under /blog/. The file is not the whole story, though: with Disallow AI Training or Block, Cloudflare also blocks training-only crawlers such as OpenAI's at its edge (Cloudflare), whatever robots.txt says. Clean up overlapping groups so the file says what you mean.

Also remember that a crawler named in any group ignores the User-agent: * group. If the synced block names a crawler, your * rules no longer apply to it.

Check what your file says now

  1. Open https://yourdomain/robots.txt and look for the # BEGIN Cloudflare Bot Preference Sync block.
  2. Run the RobotsGate site check on your domain. For every AI crawler in the registry, it shows allowed or blocked for a path, the group that matched, and the line that decided it, with groups combined as RFC 9309 requires.
  3. Check a few paths that matter to you, such as / and /blog/, with the path option.

The check reads robots.txt only. It can't see Cloudflare's edge blocks, WAF rules or page-level tags such as NOARCHIVE. See How to check which AI bots can crawl your site for what a robots.txt check can and can't tell you.