Allow AI search but block AI training in robots.txt
Several operators now publish separate tokens for "use my content to train models" and "show my site in AI search answers". That lets you allow one and block the other.
robots.txt is voluntary (RFC 9309). The descriptions below are the operators' own statements, current as of 2026-10-05. We have not independently tested how any crawler behaves.
Training tokens vs. search tokens, by operator
| Operator | Training (block) | Search / answers (allow) | Source |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | docs |
| Anthropic | ClaudeBot | Claude-SearchBot | docs |
| Perplexity | (none listed; Perplexity says PerplexityBot is not used to crawl content for AI foundation models) | PerplexityBot | docs |
| Google-Extended (Gemini training and grounding) | Googlebot (Google says Google-Extended does not affect Search inclusion or ranking) | docs | |
| Apple | Applebot-Extended | Applebot | docs |
| Amazon | Amazonbot ("may be used to train Amazon AI models") | Amzn-SearchBot ("does not crawl content for generative AI model training") | docs |
| Meta | Meta-ExternalAgent | Meta-WebIndexer | docs |
| Mistral AI | MistralAI-Training | MistralAI-Index | docs |
| DuckDuckGo | (none listed; DuckDuckGo says DuckAssistBot data is not used to train AI models) | DuckAssistBot | docs |
| You.com | (none listed; You.com describes YouBot as the crawler that powers its search engine) | YouBot | docs |
Example robots.txt
# Ask AI training crawlers not to collect content User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: Meta-ExternalAgent User-agent: Amazonbot User-agent: MistralAI-Training User-agent: CCBot Disallow: / # Allow AI search / answer crawlers (repeat any shared Disallow paths here) User-agent: OAI-SearchBot User-agent: Claude-SearchBot User-agent: PerplexityBot User-agent: Meta-WebIndexer User-agent: Amzn-SearchBot User-agent: MistralAI-Index User-agent: DuckAssistBot User-agent: YouBot User-agent: Applebot Disallow: /admin/ User-agent: * Disallow: /admin/ Sitemap: https://example.com/sitemap.xml
Why is Disallow: /admin/ listed twice? A crawler that finds a group with its own name ignores the * group, so any shared rules have to be repeated. The generator does this for you.
Caveats from the operators' own pages
- Apple: Applebot data may also be used for AI-generated answers in Siri and Search. Apple says publishers can opt out of that use with the
nosnippetrobots meta tag. robots.txt does not control it. - Meta: Meta describes Meta-ExternalAgent as covering training or "improving products by indexing content directly". Blocking it may therefore affect more than training.
- Common Crawl (CCBot) builds an open dataset that anyone can use. Whether to block it is your call; its page does not describe it as an AI-training crawler.
- Google: Google-Extended covers Gemini training and grounding in Gemini Apps and Vertex AI, so blocking it affects both.
- User-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher, Amzn-User, MistralAI-User) fetch pages when a user asks. Several operators say these may not follow robots.txt. Anthropic and Mistral describe their user agents as controllable through robots.txt.
Build this file with the generator by setting "AI training" and "Open web datasets" to Block and "AI search / answers" to Allow. Then confirm the result in the validator.