How to check which AI bots can crawl your site
Your robots.txt decides which AI crawlers are asked to stay away. Here is how crawlers read it, how to check it quickly, and what it can't tell you.
1. Look at your robots.txt
Open https://yourdomain/robots.txt. Every host has its own file, so check www., the bare domain and any subdomains separately.
- If the file returns 404 (or another 4xx), RFC 9309 treats it as "unavailable", and crawlers may access anything.
- If it returns a 5xx error or can't be reached, RFC 9309 says crawlers must assume everything is disallowed.
2. How a crawler picks its rules (RFC 9309)
- It looks for groups whose
User-agentmatches its token, ignoring case. All matching groups are combined. - If no group names it, it uses the
User-agent: *group. If there is none, everything is allowed. - Within the chosen rules, the longest matching
Allow/Disallowpath wins. When an Allow and a Disallow match equally, Allow wins.*matches any characters, and$marks the end of the URL. /robots.txtitself is always allowed.
Some operators add fallbacks. Apple says that if Applebot isn't named but Googlebot is, Applebot follows the Googlebot rules. Amazon says Amzn-SearchBot follows the directives for other search bots if it isn't named (Apple, Amazon).
3. Run an automatic check
Enter your domain in the RobotsGate site check, or call the API:
curl 'https://robotsgate.mike-tusa.workers.dev/api/check?url=https://example.com&path=/blog/'
For each AI crawler in the registry, you get allowed or blocked for that path, the group that matched (exact, documented fallback, or *), and the line that decided it. The check also flags syntax errors, misspelled tokens (for example GPT-Bot instead of GPTBot), rules placed before any User-agent line, and HTML pages served in place of robots.txt.
4. What a robots.txt check can't tell you
- Firewall/CDN rules: a WAF or bot-management setting may block or allow bots no matter what robots.txt says.
- Page-level directives: meta robots tags and
X-Robots-Tagheaders are separate. Apple, for example, documentsnosnippet, and Amazon documentsnoarchiveas "do not use the page for model training". - User-triggered fetchers such as ChatGPT-User and Perplexity-User may not follow robots.txt, according to their operators.
- Crawlers not in the registry: RobotsGate only lists tokens it could verify against official documentation.
5. Is that request really from the crawler?
Anyone can fake a user-agent string. Several operators publish ways to verify their traffic:
- OpenAI publishes IP ranges for GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI).
- Perplexity publishes IP lists for PerplexityBot and Perplexity-User (Perplexity).
- Amazon publishes IP addresses for Amazonbot, Amzn-SearchBot and Amzn-User (Amazon).
- Apple supports reverse DNS (
*.applebot.apple.com) and publishes CIDR prefixes (Apple). - Common Crawl supports reverse DNS (
*.crawl.commoncrawl.org) and publishes a JSON list of IP ranges. It also warns that some crawlers falsely claim to be CCBot (Common Crawl). - DuckDuckGo publishes DuckAssistBot IP addresses (DuckDuckGo).