RobotsGate

How to check which AI bots can crawl your site

Your robots.txt decides which AI crawlers are asked to stay away. Here is how crawlers read it, how to check it quickly, and what it can't tell you.

robots.txt is voluntary (RFC 9309). A check shows what your file asks crawlers to do. It does not show whether a given crawler obeys.

1. Look at your robots.txt

Open https://yourdomain/robots.txt. Every host has its own file, so check www., the bare domain and any subdomains separately.

2. How a crawler picks its rules (RFC 9309)

  1. It looks for groups whose User-agent matches its token, ignoring case. All matching groups are combined.
  2. If no group names it, it uses the User-agent: * group. If there is none, everything is allowed.
  3. Within the chosen rules, the longest matching Allow/Disallow path wins. When an Allow and a Disallow match equally, Allow wins. * matches any characters, and $ marks the end of the URL.
  4. /robots.txt itself is always allowed.

Some operators add fallbacks. Apple says that if Applebot isn't named but Googlebot is, Applebot follows the Googlebot rules. Amazon says Amzn-SearchBot follows the directives for other search bots if it isn't named (Apple, Amazon).

3. Run an automatic check

Enter your domain in the RobotsGate site check, or call the API:

curl 'https://robotsgate.mike-tusa.workers.dev/api/check?url=https://example.com&path=/blog/'

For each AI crawler in the registry, you get allowed or blocked for that path, the group that matched (exact, documented fallback, or *), and the line that decided it. The check also flags syntax errors, misspelled tokens (for example GPT-Bot instead of GPTBot), rules placed before any User-agent line, and HTML pages served in place of robots.txt.

4. What a robots.txt check can't tell you

5. Is that request really from the crawler?

Anyone can fake a user-agent string. Several operators publish ways to verify their traffic: