Is your robots.txt blocking AI crawlers?

Paste your robots.txt below and this page tells you, for every AI crawler that matters, whether it can read your site and which line decides. If something is closed that you meant to leave open, you also get a corrected file to copy. It runs in your browser and sends nothing anywhere.

Paste the whole file. It is read inside this page — not uploaded, not stored, not sent anywhere.

AI crawlers and what each one decides

A search robot feeds the answers an assistant gives today. A training robot feeds the next model. Blocking them costs you different things.

OAI-SearchBot OpenAI · search index
GPTBot OpenAI · model training
ChatGPT-User OpenAI · fetches a page a person asked for
Claude-SearchBot Anthropic · search index
ClaudeBot Anthropic · model training
Claude-User Anthropic · fetches a page a person asked for
anthropic-ai Anthropic · model training · retired token
PerplexityBot Perplexity · search index
Perplexity-User Perplexity · fetches a page a person asked for
Google-Extended Google · model training
GoogleOther Google · model training
Applebot Apple · search index
Applebot-Extended Apple · model training
meta-externalagent Meta · model training
FacebookBot Meta · model training · retired token
Amazonbot Amazon · search index
CCBot Common Crawl · model training
Bytespider ByteDance · model training
cohere-ai Cohere · model training · retired token

Blocking the training robot is not the same as blocking the search one

Each assistant sends more than one robot, and they do different jobs. GPTBot collects text to train future OpenAI models. OAI-SearchBot builds the index ChatGPT quotes from when somebody asks a question today. Anthropic and Perplexity split their crawlers the same way. Closing the first changes nothing you will notice this year; closing the second takes your pages out of the answers your buyers are reading right now.

That is why a "block the AI bots" snippet copied from a forum thread does more damage than its author intended. It names every token someone could find, search robots included, and the site owner walks away believing they opted out of training. The table below labels each robot with the job it does, so the decision can be made one robot at a time.

How the file is read

The text you paste goes through the same parser our audit uses on real sites: group headings, Allow winning over Disallow, longest match winning over both, wildcards and the end-of-path marker. robots.txt has more corners than it looks. Several User-agent lines before one block of rules are a single group covering all of them, and a naive reading reports the first name as allowed while it is closed.

Every robot gets a verdict and the line number behind it, so you can open your own file and see the same thing. Where a rule closes a path rather than the whole site, the path is named instead of the line. Nothing is fetched from your domain, and that is deliberate: a page that went and got the file for you would have to route the request through our servers.

What this checker is not

  • It is not a check of your server. This file states a policy; a firewall or bot-protection rule can still refuse a crawler the file welcomes, and only your access log shows that.
  • It is not a promise of a visit. An open door means a crawler may come, never that it will or that it will use what it finds.
  • It does not read your site. It reads the text in the box, so a file you edited and never deployed passes here and changes nothing out there.

Questions about robots.txt and AI crawlers

Where is my robots.txt, and how much of it should I paste?
At the root of your domain, as yoursite.com/robots.txt. Open that address in a browser and whatever you see is the file every crawler reads. Paste all of it: a rule further down can override the part you pasted, so a partial paste gives a confidently wrong answer.
Should I block GPTBot?
That is a business decision, and this page exists so you can make it knowingly. Blocking GPTBot keeps your text out of future OpenAI training runs and does not remove you from ChatGPT answers today — OAI-SearchBot is the one that does. Allowing the search robots while refusing the training ones is a coherent position, and plenty of publishers hold it.
Does robots.txt actually stop anyone?
It stops the ones that read it, which includes the named operators here. It is a request rather than a fence: nothing enforces it, and a scraper that ignores it simply takes the page. If you need the door physically shut, that is a server rule, not a text file.
Why is a crawler I never mentioned shown as blocked?
Because a User-agent: * group applies to every robot without a group of its own. One Disallow: / line there closes the site to all of them at once, which is the most common way a site drops out of AI answers with nobody having decided to.