GPTBot, ClaudeBot, PerplexityBot: The Five Names in Your robots.txt Are Not One Thing

Five names, one mental model, and the mental model is wrong

Somewhere out there is a robots.txt file that got pasted in three years ago from a blog post titled something like "Block AI From Stealing Your Content." It has five or six User-agent lines in it, a Disallow: / under each, and nobody has looked at it since. The business owner who approved it thinks they're protected. They have no idea what they actually did.

That's the failure mode our AI Crawler Access module exists to catch, and it's the one I see catch people most often, by accident rather than on purpose. Five names show up in that module: GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, and Bingbot. Looked at from a distance, they read like five flavors of the same thing, five bots doing the same job for five companies. They are not. OpenAI runs three separate crawlers under three separate names, and has since 2024. Anthropic does the same. Bingbot deliberately doesn't split at all, and blocking it costs you more than you'd guess. And at least one of the five has been accused, by name, of not honoring robots.txt in the first place.

None of that is brand new information, exactly. The bots have existed in their current form for a while. What's new is that the companies running them have finally started saying, in writing, what blocking each one actually costs you, and that documentation only started catching up within the last year. If you set your robots.txt file once and haven't looked at it since, you're working from an old map of a landscape that quietly got more detailed.

The field guide nobody handed you

OpenAI runs three bots, not one. This has been true since 2024, which makes it old news to anyone paying close attention and news to almost everyone else.

  • GPTBot trains OpenAI's foundation models. Disallowing it opts you out of training use. This is the one most "block all AI" snippets are actually aiming at, whether or not they say so.
  • OAI-SearchBot fetches content specifically for ChatGPT's live search and citation feature. Disallow this one and you will not appear in ChatGPT's search answers. This is a different consequence from opting out of training, and it's the one that actually costs a small business visibility today.
  • ChatGPT-User isn't a crawler in the automated sense at all. It fires when an actual person, mid-conversation, asks ChatGPT to go look at a specific page right now. OpenAI's crawler documentation was revised in what's reported as December 2025, and the revision loosened the language around whether ChatGPT-User obeys robots.txt the way a crawler would. It's not that the rules stopped applying outright. It's that OpenAI's own docs no longer promise this particular identity behaves like GPTBot or OAI-SearchBot when it comes to respecting your Disallow lines.

Three names, three purposes, three different things you're actually deciding when you block one of them.

Anthropic runs three bots too, and this part is where the framing matters. ClaudeBot, Claude-SearchBot, and Claude-User existed before the news that prompted this post. They didn't just get invented. What happened, spotted by researcher Pedro Dias and reported by Barry Schwartz at Search Engine Roundtable in late February, is that Anthropic updated its crawler documentation to spell out, for the first time, what disabling each of the three actually does to your site's presence with Claude. ClaudeBot trains models. Claude-SearchBot fetches for citation and search-style answers. Claude-User fires on behalf of a live user, the same shape as ChatGPT-User. The bots aren't new. The clarity is. Until that update, a business owner blocking ClaudeBot had no official document telling them whether that also touched Claude's citation surface or its live-user fetches. Now they do, and it's worth going and reading it if you've touched your robots.txt in that window.

PerplexityBot is the one on this list with an asterisk attached, and it gets its own section below, because the asterisk is bigger than "it's just another crawler."

Bingbot deliberately doesn't split, and that's the interesting design choice on this whole list. One crawler identity feeds classic Bing search results, Yahoo Search, DuckDuckGo, and Microsoft Copilot's grounded answers. There is no Bing equivalent of Google's Google-Extended, no separate identity you can allow for search while blocking for AI. Block Bingbot and you've blocked all four of those surfaces in the same stroke, whether that was the intent or not. Microsoft published guidance around early May 2026 explaining that the index built for AI grounding differs internally from the ranking index used for classic search results, in terms of factual fidelity, source attribution, freshness, and coverage, even though the same crawler populates both. Same bot, two different downstream uses, no lever to separate them at the robots.txt level. If you want Bing search traffic at all, Bingbot has to be let in, full stop, and Copilot rides along whether you wanted that bundled or not.

What actually changed recently, and when

Here's where I want to slow down, because the headline version of this story ("Anthropic split ClaudeBot into three bots") is wrong, and the correction is the actual story.

Anthropic's documentation update landed around February 20, 2026, and it did not create Claude-SearchBot or Claude-User. Those identities already existed. What the update did was make explicit, bot by bot, what happens when you disable each one, information that simply wasn't written down anywhere official before that date. If you were managing a robots.txt file with an opinion about Claude before late February, you were forming that opinion without the documentation that now exists to inform it.

OpenAI's change is a separate event, reported around December 2025, and it moved in a slightly different direction. Rather than adding clarity about consequences, OpenAI loosened the compliance language specifically around ChatGPT-User's relationship to robots.txt. The practical read: don't assume ChatGPT-User treats your Disallow rules with the same weight GPTBot does.

Put those two together and the real "why now" isn't that the crawlers changed. It's that the paperwork did, in both directions, within a few months of each other, after most of the internet's robots.txt files were already written by someone assuming none of this would ever change.

What publishers who bother to specify are actually doing

A site called technologychecker.io ran an analysis of Cloudflare Radar's public robots.txt data: not a Cloudflare publication itself, and not traffic data. It's a count of how many published robots.txt files mention each bot by name, and under which directive, across a single-day snapshot, dated August 31, 2026, of a few thousand published files from Cloudflare's sample of top domains. That's a small slice of the web, and the numbers only describe sites that bothered to name these bots at all, not the internet at large.

Here's what the snapshot found:

Bot Disallow mentions Allow mentions Ratio
GPTBot 696 299 ~2.3 : 1
ClaudeBot 619 259 ~2.4 : 1
OAI-SearchBot 250 266 roughly even, slightly more allow
Bingbot 200 195 roughly even

That table leaves out PerplexityBot entirely. Not by choice on my part; the source snapshot doesn't break out a separate row for it, so I'm not putting a number there that isn't in the data.

Read that as: among publishers who bother to name GPTBot specifically, they block it roughly 2.3 times for every time they allow it. Same shape for ClaudeBot, at roughly 2.4 to 1. That's a real signal about the training-bot identities specifically. But notice what happens with OAI-SearchBot, the bot that actually controls whether you show up in ChatGPT's citations: the ratio flips almost dead even, tilted slightly toward allow. And Bingbot, the one bot that can't be selectively blocked without losing Bing, Yahoo, and DuckDuckGo along with Copilot, sits close to a coin flip too.

That pattern is consistent with what the field guide above would predict. Training bots get blocked more often, because training is the use case people actually have an opinion about. Search and citation bots get treated more cautiously, because blocking them has an obvious, immediate cost that training-bot blocking doesn't. Whether that's informed strategy or accidental luck on the part of the sites in this sample, I can't tell you from here. It's a single-day snapshot of naming behavior across a few thousand files, not a study of intent.

The uncomfortable exception

Here's the part of this list that doesn't resolve as neatly, and it's old news, not a fresh discovery.

On August 4, 2025, Cloudflare published a blog post accusing Perplexity of ignoring robots.txt outright. The claim, specifically: Cloudflare documented not just the declared PerplexityBot crawler but an undeclared "stealth" variant, one that impersonated Chrome running on macOS and rotated IPs and ASNs, apparently to keep crawling sites that had already blocked the declared bot. Cloudflare responded by de-listing Perplexity from its verified-bot program. Perplexity's spokesperson denied the accusation and called it a publicity stunt.

That was thirteen months ago as of this writing, and Cloudflare has never publicly announced restoring Perplexity's verified-bot status since. That makes this a long-running, unresolved dispute between two companies with obvious competing interests, not breaking news. Cloudflare has a business reason to police bot traffic aggressively. Perplexity has a business reason to deny scraping accusations. What's true regardless of who you believe is this: a green checkmark next to PerplexityBot in a robots.txt-based tool, ours included, tells you the file asked correctly. It does not, and cannot, tell you whether every crawler that reads that file actually honors what it says. That gap is real and it's worth knowing about, not something a scoring tool can close from the outside.

Why "allow" isn't automatically free

There's a version of this whole conversation that assumes blocking is bad and allowing is good, full stop. The 2024 record on ClaudeBot specifically says otherwise.

In July 2024, iFixit's CEO reported that ClaudeBot had hit their servers roughly a million times in 24 hours. Around the same time, Freelancer.com's CEO reported millions of ClaudeBot requests arriving within a matter of hours. Neither of those sites was making a philosophical point about AI training. They were watching a crawler behave like a denial-of-service event.

That's useful context for why granular, per-bot control became something publishers actually wanted rather than an abstract policy debate. "Allow every AI bot" sounds like the generous, forward-thinking choice, right up until one of them decides your server can handle a million requests a day. This predates the documentation changes discussed above, and it's a big part of why those changes matter now: once you know a bot's specific purpose, you can make a deliberate choice about it instead of an all-or-nothing gamble.

The accidental-blocking problem

A small study gives a useful, honestly-scoped look at how this actually plays out on real sites. Kate Shaw at Relevance checked robots.txt files on 50 sites across 10 industries in August 2026, looking at GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and CCBot.

Eleven of the 50 sites, 22%, blocked at least one of those five. GPTBot and CCBot were blocked most often, 9 of 50 each, 18%. ClaudeBot was blocked on 7 of 50, 14%. Google-Extended on 8%. PerplexityBot was blocked least of all, on just 2 of 50, 4%.

That last number stopped me for a second, next to the previous section. The one bot with a documented, still-open accusation of ignoring robots.txt outright is the one publishers in this sample bothered to block least often. That's not damning of anyone in the study. It's a small sample, and there's no way to know whether those particular sites made an informed choice or just never got around to naming Perplexity specifically. But it's an irony I can't shake: the crawler with the shakiest compliance record is the one getting the least explicit attention.

News and media sites stood out from the rest: 5 of 7 blocked at least one of the five bots, a real difference worth noting even though 7 sites is too small a base to turn into a clean percentage. The study itself is upfront about what it doesn't cover, too. It looked at robots.txt files specifically, and disclaims edge- or WAF-level blocking as outside its method entirely, which is the same gap our own module has, discussed below.

What to actually check on your own site

Open your robots.txt file. It usually lives at yourdomain.com/robots.txt. Look for these exact strings, not close approximations: GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Bingbot. If you see one of them followed by Disallow: / with nothing after the slash, that identity is blocked at the root, completely, for everything. If instead you see something narrower, like Disallow: /admin/ or Disallow: /cart/, that's a partial block on a specific section, and it's a different, much less consequential thing than blocking the whole site.

While you're in there, look for User-agent: * followed by Disallow: /, with no bot-specific line underneath it that overrides that block. That wildcard rule catches everything not named individually, including every one of these five bots, and it's the single most common way a site ends up blocking AI crawlers without anyone deciding to.

Here's what our own AI Crawler Access module actually does with what it finds. It's worth 15 of the report's 100 points, split evenly across the five named bots above: 3 points each for a bot that is not blocked at the root. "Blocked at the root" means a specific Disallow: / for that bot, or the wildcard catch-all with no override, not a narrower path block, which the module treats as allowed and doesn't surface as a separate finding at all. There's also a separate check for a noindex meta tag, reported as its own PASS or FAIL finding, but it carries zero points in this module today. That's a known gap in the current scoring, not something the module's score covers.

The tool also knows the difference between a robots.txt file you can edit and one you can't. When it detects that the file is managed by Cloudflare rather than hand-written, it rewrites its guidance to say so, because telling someone to go edit a file their CDN controls isn't useful advice.

What it genuinely can't see is everything this post has been circling. Edge-level and WAF bot-management rules that never touch robots.txt at all. Rate-limiting that throttles a crawler without ever declaring it blocked. And the exact kind of stealth-crawling Cloudflare has accused Perplexity of, a bot that never announces itself and therefore never shows up as a line in anybody's robots.txt file, ours included. It also doesn't separately check Claude-SearchBot, Claude-User, or ChatGPT-User by name. Name those three yourself anyway, even without a scoring engine asking you to, because they're the identities most likely to matter to your actual citation surface going forward, and the documentation now exists to tell you what each one costs if you get it wrong. The free Grade Check runs the same underlying analyzer as an instant-fail gate on this exact question, checking whether any of the five named bots is fully blocked at the root.

None of that makes the check worthless, and I don't think it should read that way. It means a passing score on this module tells you something specific and true: your site is asking correctly. Robots.txt is a published request, not a lock on a door. What it can't promise is that every crawler reading that request is going to honor it, and as of this writing, at least one of the five names in the checklist has an open, unresolved question mark next to whether it does.

If you want to know where your own site actually stands on this today, not from a blog post you pasted three years ago but from an actual read of your current file, run the free Grade Check. It takes about a minute, costs nothing, and it'll tell you plainly which of these five names, if any, your site is quietly turning away.