Matching rules (RFC 9309)
- A crawler obeys the one group whose
User-agentmatches its product token. It falls back to*only if no named group matches. So aGPTBotgroup completely replaces the*rules for GPTBot. - Within a group, the longest matching path wins. If an Allow and a Disallow are equally long, Allow wins.
*matches any sequence of characters and$anchors the end of the URL. Matching is case-sensitive for paths.Crawl-delayis ignored by Google. robots.txt controls crawling, not indexing: a blocked URL can still appear in results if other pages link to it.
AI crawlers
Several AI companies use separate tokens for training and for live retrieval. GPTBot collects training data while OAI-SearchBot and ChatGPT-User fetch pages for answers. Google-Extended controls use in Gemini and Vertex AI without affecting Google Search. ClaudeBot handles training and Claude-User handles user-initiated fetches. Blocking one does not block the others. Respecting robots.txt is voluntary, so the file is a request, not an access control.
FAQ
Does this fetch my live robots.txt?
No, paste the file content. That keeps the tool private and lets you test changes before deploying them.
Why is Googlebot allowed when * is blocked?
If a group names Googlebot, Googlebot follows only that group and ignores the * rules.