How to block GPTBot and other AI crawlers in robots.txt
Add a group for each crawler's robots.txt token with Disallow: /. This page lists the tokens that operators document for AI training and explains the details that often trip people up.
1. Block GPTBot (OpenAI)
OpenAI says GPTBot "is used to crawl content that may be used in training our generative AI foundation models" and that disallowing it "indicates a site's content should not be used in training" (OpenAI: Overview of OpenAI Crawlers).
User-agent: GPTBot Disallow: /
This does not block OAI-SearchBot, which OpenAI uses to show sites in ChatGPT search results. OpenAI says each setting is independent.
2. Block AI training and open-dataset crawlers
You can list several User-agent lines above one set of rules. Every token below except CCBot is described by its operator as related to AI training. CCBot is Common Crawl's crawler for an open web dataset; Common Crawl's page does not describe it as an AI-training crawler, so include it only if you also want to stay out of open datasets.
User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: Meta-ExternalAgent User-agent: Amazonbot User-agent: MistralAI-Training User-agent: CCBot Disallow: /
| Token | Operator's description (summarised) | Source |
|---|---|---|
| ClaudeBot | Collects web content that could contribute to Anthropic model training | Anthropic |
| Google-Extended | A control token, not a separate crawler. It manages whether content Google crawls may be used to train Gemini models and for grounding in Gemini Apps / Vertex AI. Google says it does not affect inclusion in Google Search. | |
| Applebot-Extended | Opts content out of training Apple's foundation models. Apple says it "does not crawl webpages", and pages that disallow it can still appear in Apple search results. | Apple |
| Meta-ExternalAgent | Crawls for uses "such as training foundation AI models or improving products by indexing content directly" | Meta |
| Amazonbot | Used to improve Amazon products and services, and "may be used to train Amazon AI models" | Amazon |
| MistralAI-Training | Crawls content to build training datasets for Mistral models | Mistral AI |
| CCBot | Open-dataset crawler (not described as AI training by Common Crawl). It builds an open repository of web crawl data that anyone can use. | Common Crawl |
The full registry has more tokens, and it also lists the ones we could not verify. One example is Bytespider: we could not load an official ByteDance page to confirm it, so we leave it out of generated rules. You can still add it yourself as a custom rule.
3. Mistakes to avoid
- A named group replaces
*. Under RFC 9309, a crawler uses the group that names it and ignores theUser-agent: *group. If you addUser-agent: GPTBotwith onlyAllow: /, your*disallow rules no longer apply to GPTBot. - Use the token, not the full user-agent string. Write
User-agent: GPTBot, notGPTBot/1.4or the whole browser-like string. - Every host needs its own file. Anthropic and Amazon both say robots.txt is read per host or subdomain. So
blog.example.comneeds its own/robots.txt. - Changes take time. OpenAI mentions about 24 hours for search changes. Perplexity says up to 24 hours. Meta says crawlers may cache robots.txt for up to 24 hours. DuckDuckGo says 72 hours for DuckAssistBot. Amazon says it may use a cached copy up to 30 days old.
- Blocking IP addresses can backfire. Anthropic and Cohere both note that IP blocking can stop their crawler from reading your robots.txt, so the opt-out may not be recorded.
- Your file is public. RFC 9309 points out that listing paths in robots.txt exposes them to anyone.
4. User-triggered fetchers are different
Some agents fetch a page only when a person asks an AI assistant about it. OpenAI says robots.txt rules "may not apply" to ChatGPT-User. Perplexity says Perplexity-User "generally ignores robots.txt rules". Meta says Meta-ExternalFetcher "may bypass robots.txt". Amazon says Amzn-User "may not follow all robots.txt directives". You can still list them, but don't count on robots.txt to stop them.
5. Check your result
Paste your file into the validator, or run the site check on your live domain. You'll see which AI crawlers are allowed or blocked and which line decides each one.