RobotsGate

Allow AI search but block AI training in robots.txt

Several operators now publish separate tokens for "use my content to train models" and "show my site in AI search answers". That lets you allow one and block the other.

robots.txt is voluntary (RFC 9309). The descriptions below are the operators' own statements, current as of 2026-10-05. We have not independently tested how any crawler behaves.

Training tokens vs. search tokens, by operator

OperatorTraining (block)Search / answers (allow)Source
OpenAIGPTBotOAI-SearchBotdocs
AnthropicClaudeBotClaude-SearchBotdocs
Perplexity(none listed; Perplexity says PerplexityBot is not used to crawl content for AI foundation models)PerplexityBotdocs
GoogleGoogle-Extended (Gemini training and grounding)Googlebot (Google says Google-Extended does not affect Search inclusion or ranking)docs
AppleApplebot-ExtendedApplebotdocs
AmazonAmazonbot ("may be used to train Amazon AI models")Amzn-SearchBot ("does not crawl content for generative AI model training")docs
MetaMeta-ExternalAgentMeta-WebIndexerdocs
Mistral AIMistralAI-TrainingMistralAI-Indexdocs
DuckDuckGo(none listed; DuckDuckGo says DuckAssistBot data is not used to train AI models)DuckAssistBotdocs
You.com(none listed; You.com describes YouBot as the crawler that powers its search engine)YouBotdocs

Example robots.txt

# Ask AI training crawlers not to collect content
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Amazonbot
User-agent: MistralAI-Training
User-agent: CCBot
Disallow: /

# Allow AI search / answer crawlers (repeat any shared Disallow paths here)
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Meta-WebIndexer
User-agent: Amzn-SearchBot
User-agent: MistralAI-Index
User-agent: DuckAssistBot
User-agent: YouBot
User-agent: Applebot
Disallow: /admin/

User-agent: *
Disallow: /admin/

Sitemap: https://example.com/sitemap.xml

Why is Disallow: /admin/ listed twice? A crawler that finds a group with its own name ignores the * group, so any shared rules have to be repeated. The generator does this for you.

Caveats from the operators' own pages

Build this file with the generator by setting "AI training" and "Open web datasets" to Block and "AI search / answers" to Allow. Then confirm the result in the validator.