API reference
A JSON API. No key is needed for the free tier. All endpoints send Access-Control-Allow-Origin: *, so you can call them from a browser. Optional Pro license keys raise rate limits (see Pro).
Errors
Errors come back as JSON with a matching HTTP status:
{ "error": { "code": "invalid_input", "message": "Invalid generate request.", "details": ["categories.training must be one of allow, block, omit"] } }
| Status | code | When |
|---|---|---|
| 400 | invalid_input, invalid_json, url_not_allowed | Bad parameters, bad JSON, or a URL that is not allowed (non-http(s), private/local address, credentials, non-default port) |
| 404 | not_found | Unknown endpoint |
| 405 | method_not_allowed | Wrong HTTP method |
| 413 | payload_too_large | Request body over 1,000,000 bytes, or robots_txt over 512,000 bytes |
| 429 | rate_limited | Too many requests from one client. Free: 20 /api/check or 60 /api/generate+/api/validate per 60 seconds. Pro (valid license): 300 check or 600 generate+validate per 60 seconds. Response includes Retry-After. |
| 415 | unsupported_media_type | POST without Content-Type: application/json |
| 422 | too_complex | The robots.txt needs more work than the per-request limit allows (/api/validate and /api/check). The message and details ({stage, limit}) say which stage ran out (parse, normalize, compile or match); on /api/check it says the site's robots.txt is too large or complex to analyse, and when the match stage ran out it suggests checking a shorter path. Practical limits: about 34,000 plain rules (the whole 512,000-byte input), about 18,000 rules with a literal prefix or $ suffix, about 12,000 rules starting with /* (on a 2,048-character path, about 4,000 when the rule text is absent from the path and about 700 when it nearly matches everywhere), and about 70,000 %-escapes in one rule. This is a limit of this checker, not of robots.txt: Google reads up to 500 KiB. |
| 502 / 504 | upstream_error, too_many_redirects, dns_check_failed / upstream_timeout | /api/check could not fetch the site's robots.txt (timeout is 8 seconds), or the DNS safety check could not be completed. The check fails closed, so if the DNS lookup errors the request is refused. |
GET /api/crawlers
Returns the AI crawler registry. Optional ?category=training|dataset|search|user_fetch|other.
Each entry has: name (the robots.txt token), operator, category, purpose, respects_robots (yes, no_user_triggered, or not_stated, based on what the operator documents), robots_notes, docs_url (the operator's page), date_checked, and verification_method. The response also lists unverified tokens that we could not confirm from official documentation, and these are never used in generated rules.
personal_agents (added in 0.2.7) covers user-driven browsing agents such as Meta's Muse, which are not crawlers. It has description, date_checked, robots_txt_controls_them (not_reliably), agents, options (allow, rate-limit, challenge, require login, block, watch for the standard, each with its trade-off), signed_agents_note, sources and disclaimer. Each agent shows what its operator has published (operator_published: user_agent, ip_ranges, robots_txt_token, request_signature; null means nothing is published) and, separately, third_party_listings such as Cloudflare Radar's (with verified_by_cloudflare, Cloudflare's own label). As of 2026-10-08 Meta publishes none of these for Muse, so Muse is not a registry crawler and is never used in generated rules. See Should I let Muse agents onto my site? RobotsGate is not affiliated with Meta.
curl https://robotsgate.mike-tusa.workers.dev/api/crawlers?category=training
POST /api/generate
Builds a robots.txt file.
| Field | Type | Meaning |
|---|---|---|
| categories | object | Map of category to allow, block or omit (omit means no rule, so the * group applies) |
| agents | object | Per-token overrides, e.g. {"GPTBot":"block"}. Tokens not in the registry are accepted (letters, digits, _ . -, up to 64 characters) and the response includes a note about them. Max 100. |
| default | "allow" | "block" | Policy for User-agent: * (default allow) |
| disallow_paths / allow_paths | string[] | Paths for *. These are also copied into the group of explicitly allowed AI crawlers, because a named group replaces the * group. Max 100 paths, 512 characters each. |
| sitemaps (or sitemap) | string[] (string) | Absolute http(s) sitemap URLs, max 20 |
| custom_rules | string | Appended as written (max 20,000 bytes). Any parse problems are returned in custom_rules_diagnostics. |
curl -X POST https://robotsgate.mike-tusa.workers.dev/api/generate \
-H 'content-type: application/json' \
-d '{"categories":{"training":"block","dataset":"block","search":"allow"},"disallow_paths":["/admin/"],"sitemaps":["https://example.com/sitemap.xml"]}'
Response: { robots_txt, summary: [{name, action, source, in_registry}], notes: [], custom_rules_diagnostics, personal_agents_note }. personal_agents_note (0.2.7) is a reminder that robots.txt can't target personal agents such as Meta's Muse. RobotsGate is not affiliated with Meta.
POST /api/validate
Body: { "robots_txt": "...", "path": "/blog/post" } (path is optional and defaults to /).
Response: valid, errors, warnings, info (each item has {line, code, message}), stats, and ai_crawlers. Each ai_crawlers item has status (allowed/blocked for the path), root_blocked, matched_group, match_type (exact, fallback, wildcard, none), and deciding_rule.
personal_agents (added in 0.2.7, also in /api/check results; null when robots.txt was unreachable) explains what the file can and can't do about personal agents such as Meta's Muse. It never changes valid, errors, warnings or ai_crawlers. Fields: robots_txt_controls_them (not_reliably), declared_ai_crawlers_blocked (count for the path), agents_without_published_identifier, findings (each {code, level, message}, level notice or info), guide_url and as_of. Finding codes: declared_crawlers_only (the file blocks declared AI crawlers, which doesn't reliably reach personal agents), wildcard_not_agent_block (User-agent: * / Disallow: / is not a block for personal agents), meta_crawlers_not_muse (blocking Meta's crawlers doesn't cover Muse), muse_token_unpublished (the file names a Muse token Meta hasn't published), and no_ai_blocks. The object also carries disclaimer ("RobotsGate is not affiliated with Meta.").
deciding_rule is {type, path, line} or null; a path longer than 256 characters is cut to its first 256 and the rule gets path_truncated: true.
diagnostics (also in /api/check results): {truncated, limit, errors, warnings, info} with the total counts. Each of errors, warnings and info lists at most limit (500) entries; truncated is true when any list was cut.
Matching follows RFC 9309. Groups are matched by product token without regard to case, and groups for the same token are combined. If no group names the crawler, the * group applies. The longest matching rule wins, and when an Allow and a Disallow rule are equally long, Allow wins. * and $ wildcards are supported. The fallback match type is used where the operator documents one: Apple and ImageSift say their bots follow Googlebot rules when they are not named.
GET /api/check?url=https://example.com&path=/
Fetches /robots.txt from the site's origin and returns the same analysis as /api/validate, plus http_status, final_url, redirects and robots_found.
- Only
http/httpson the default ports. Credentials in URLs, localhost,.local/.internal-style names, and private, loopback, link-local or reserved IP addresses are rejected. Hostnames are also resolved over DNS-over-HTTPS, and requests are rejected if the name points to a private address. If that DNS check cannot be completed (lookup error, timeout, or DNS server failure), the request is refused with 502dns_check_failed. - Up to 5 redirects. Each hop is checked again.
- 8-second timeout. Only the first 512,000 bytes are analysed.
- Rate limit (free): 20 requests per 60 seconds per client IP (an IPv6 client counts per /64). Pro: 300 per 60 seconds per license. Over the limit you get 429
rate_limitedwith aRetry-Afterheader. Free limits are counted separately in each Cloudflare location and are approximate. - Caching: results for each site are cached for up to 5 minutes. The response includes
cache.hitand anX-RobotsGate-Cacheheader, andfetched_atshows when robots.txt was actually fetched. Thepathis evaluated fresh on every request. - The checker identifies itself as
RobotsGate/0.2.9 (+https://robotsgate.mike-tusa.workers.dev; robots.txt checker; fetches /robots.txt only). - A 4xx response is reported as "unavailable": under RFC 9309, crawlers may access anything. A 5xx response is reported as "unreachable": under RFC 9309, crawlers should assume complete disallow.
MCP server, llms.txt and OpenAPI
RobotsGate is also a remote MCP server at POST https://robotsgate.mike-tusa.workers.dev/mcp (Streamable HTTP, JSON-RPC 2.0, stateless, JSON responses, no sessions; GET and DELETE return 405). Protocol versions 2026-07-28 (per-request _meta) and 2025-11-25, 2025-06-18, 2025-03-26, 2024-11-05 (initialize). Most MCP clients accept this config; some use a different format (for example, VS Code uses a servers key):
{ "mcpServers": { "robotsgate": { "type": "http", "url": "https://robotsgate.mike-tusa.workers.dev/mcp" } } }
| Tool | Arguments | Same engine as | Counts against |
|---|---|---|---|
list_crawlers | category (optional) | GET /api/crawlers | The light MCP cap (about 120 per minute per IP) |
generate_robots | The /api/generate body fields | POST /api/generate | The shared generate + validate limit |
validate_robots | robots_txt (up to 50,000 characters over MCP; the whole request must fit in 64 KB), path | POST /api/validate | The shared generate + validate limit |
check_site | url, path | GET /api/check (same URL and private-address checks) | The /api/check limit |
All four tools are read-only. MCP uses the same license headers, soft-fail and counters as the JSON API (free: 20 checks and 60 generate/validate calls per 60 seconds per client IP; Pro: 300 and 600 per license). initialize, tools/list, ping and list_crawlers share a light cap of about 120 per minute per IP (HTTP 429 with Retry-After over it). Input errors and rate limits inside tools/call come back as a tool result with isError: true, _meta["robotsgate/error"] (code, httpStatus, retryAfterSeconds) and a Retry-After header. Requests are capped at 64 KB, with one tools/call per HTTP request. An Origin header must be an https origin or http loopback (otherwise 403). There is no 401: a bad key just gets the free limits.
Machine-readable docs: /llms.txt, /openapi.json (OpenAPI 3.1, generated from the code) and /.well-known/mcp/server-card.json (MCP server card). GET /api/health returns { "ok": true, "version": "…" }.
curl -s -X POST 'https://robotsgate.mike-tusa.workers.dev/mcp' -H 'content-type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"check_site","arguments":{"url":"example.com"}}}'
RobotsGate Pro
Optional paid tier ($9/mo) for higher API rate limits. Product slug robotsgate-pro; license keys use prefix RBTG.
| Free | Pro | |
|---|---|---|
/api/check | 20 / 60 s per client IP | 300 / 60 s per license |
/api/generate + /api/validate | 60 / 60 s per client IP | 600 / 60 s per license |
Send the license key on gated endpoints (/api/check, /api/generate, /api/validate) using either header:
Authorization: Bearer RBTG-… # or X-License-Key: RBTG-…
CORS allows authorization and x-license-key. A successful Pro request includes x-robotsgate-plan: pro (and x-ratelimit-remaining when available).
Soft-fail: a missing, invalid, revoked, or wrong-product key does not return 401. The request continues on the free-tier rate limits and without the Pro plan header. /api and /api/crawlers stay ungated.
curl -H 'Authorization: Bearer RBTG-…' \ 'https://robotsgate.mike-tusa.workers.dev/api/check?url=https://example.com'