RobotsGate

How to block GPTBot and other AI crawlers in robots.txt

Add a group for each crawler's robots.txt token with Disallow: /. This page lists the tokens that operators document for AI training and explains the details that often trip people up.

What robots.txt can and can't do: robots.txt is a voluntary protocol (RFC 9309). Crawlers that follow it will stay out. Nothing forces any crawler to obey it. To actually prevent access, use authentication or server/CDN-level blocking.

1. Block GPTBot (OpenAI)

OpenAI says GPTBot "is used to crawl content that may be used in training our generative AI foundation models" and that disallowing it "indicates a site's content should not be used in training" (OpenAI: Overview of OpenAI Crawlers).

User-agent: GPTBot
Disallow: /

This does not block OAI-SearchBot, which OpenAI uses to show sites in ChatGPT search results. OpenAI says each setting is independent.

2. Block AI training and open-dataset crawlers

You can list several User-agent lines above one set of rules. Every token below except CCBot is described by its operator as related to AI training. CCBot is Common Crawl's crawler for an open web dataset; Common Crawl's page does not describe it as an AI-training crawler, so include it only if you also want to stay out of open datasets.

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Amazonbot
User-agent: MistralAI-Training
User-agent: CCBot
Disallow: /
TokenOperator's description (summarised)Source
ClaudeBotCollects web content that could contribute to Anthropic model trainingAnthropic
Google-ExtendedA control token, not a separate crawler. It manages whether content Google crawls may be used to train Gemini models and for grounding in Gemini Apps / Vertex AI. Google says it does not affect inclusion in Google Search.Google
Applebot-ExtendedOpts content out of training Apple's foundation models. Apple says it "does not crawl webpages", and pages that disallow it can still appear in Apple search results.Apple
Meta-ExternalAgentCrawls for uses "such as training foundation AI models or improving products by indexing content directly"Meta
AmazonbotUsed to improve Amazon products and services, and "may be used to train Amazon AI models"Amazon
MistralAI-TrainingCrawls content to build training datasets for Mistral modelsMistral AI
CCBotOpen-dataset crawler (not described as AI training by Common Crawl). It builds an open repository of web crawl data that anyone can use.Common Crawl

The full registry has more tokens, and it also lists the ones we could not verify. One example is Bytespider: we could not load an official ByteDance page to confirm it, so we leave it out of generated rules. You can still add it yourself as a custom rule.

3. Mistakes to avoid

4. User-triggered fetchers are different

Some agents fetch a page only when a person asks an AI assistant about it. OpenAI says robots.txt rules "may not apply" to ChatGPT-User. Perplexity says Perplexity-User "generally ignores robots.txt rules". Meta says Meta-ExternalFetcher "may bypass robots.txt". Amazon says Amzn-User "may not follow all robots.txt directives". You can still list them, but don't count on robots.txt to stop them.

5. Check your result

Paste your file into the validator, or run the site check on your live domain. You'll see which AI crawlers are allowed or blocked and which line decides each one.