ParseJar › Web

robots.txt Tester

Paste a robots.txt file and a URL to see whether a crawler may fetch it, which rule decides, and how 20 common search and AI crawlers are treated.

Search & AI crawler matrix

Matching rules (RFC 9309)

  • A crawler obeys the one group whose User-agent matches its product token. It falls back to * only if no named group matches. So a GPTBot group completely replaces the * rules for GPTBot.
  • Within a group, the longest matching path wins. If an Allow and a Disallow are equally long, Allow wins.
  • * matches any sequence of characters and $ anchors the end of the URL. Matching is case-sensitive for paths.
  • Crawl-delay is ignored by Google. robots.txt controls crawling, not indexing: a blocked URL can still appear in results if other pages link to it.

AI crawlers

Several AI companies use separate tokens for training and for live retrieval. GPTBot collects training data while OAI-SearchBot and ChatGPT-User fetch pages for answers. Google-Extended controls use in Gemini and Vertex AI without affecting Google Search. ClaudeBot handles training and Claude-User handles user-initiated fetches. Blocking one does not block the others. Respecting robots.txt is voluntary, so the file is a request, not an access control.

FAQ

Does this fetch my live robots.txt?

No, paste the file content. That keeps the tool private and lets you test changes before deploying them.

Why is Googlebot allowed when * is blocked?

If a group names Googlebot, Googlebot follows only that group and ignores the * rules.

Want an AI agent to read pages politely? ScrapeMole Reader turns a URL into clean Markdown, links or a sitemap list, pay-per-call over x402. Disclosure: same operator as ParseJar.

Related tools