ClaudeBot vs Claude-SearchBot vs Claude-User: what each crawler does
ClaudeBot collects training data, Claude-SearchBot indexes pages for Claude's search and Claude-User fetches pages users ask for. Each obeys its own group.
LLaunchScaler·Published ·8 min read
Claude-SearchBot is the Anthropic crawler that indexes web pages for Claude's search results; ClaudeBot is the one that collects content for training, and Claude-User fetches a page when someone asks Claude a question that needs it. Each follows its own robots.txt group, so blocking ClaudeBot does not block the other two, and blocking all three takes your own pages out of Claude's answers.
Everything below comes from Anthropic's help article on its crawlers, checked in September 2026, and from the robots.txt standard, RFC 9309.
What is Claude-SearchBot?
Claude-SearchBot is the crawler behind Claude's web search results. Anthropic says it "analyzes online content specifically to enhance the relevance and accuracy of search responses." Blocking it "prevents our system from indexing your content for search optimization, which may reduce your site's visibility and accuracy in user search results."
In practice, it is the Anthropic crawler that decides whether Claude can find your pages when a user turns web search on. It is the Anthropic equivalent of OpenAI's OAI-SearchBot and Perplexity's PerplexityBot: a crawler whose job is citations, not training. If you want Claude to recommend your product with a link, this is the bot to keep open.
How do ClaudeBot, Claude-SearchBot and Claude-User differ?
They differ by purpose, and Anthropic describes the cost of blocking each one separately. ClaudeBot feeds training, Claude-SearchBot feeds search and Claude-User acts for a user in real time. The table uses Anthropic's own wording.
ClaudeBot
Claude-SearchBot
Claude-User
Purpose, per Anthropic
Collects web content "that could potentially contribute to their training"
Questions, answered
What people ask about this
01
What is Claude-SearchBot?
Claude-SearchBot is Anthropic's crawler for Claude's search. Anthropic says it analyses online content to improve the relevance and accuracy of search responses, and that blocking it may reduce your site's visibility and accuracy in Claude's search results.
02
What is the difference between ClaudeBot and Claude-SearchBot?
Analyses online content "to enhance the relevance and accuracy of search responses"
When "individuals ask questions to Claude, it may access websites using a Claude-User agent"
Runs
Automatically
Automatically
When a user's question calls for a page
robots.txt token
ClaudeBot
Claude-SearchBot
Claude-User
What blocking it does
Your "future materials should be excluded from our AI model training datasets"
"May reduce your site's visibility and accuracy in user search results"
"Prevents our system from retrieving your content in response to a user query"
What you keep if you block only this one
Search citations and user fetches
Training inclusion and user fetches
Training inclusion and search indexing
Anthropic's help article names the tokens but does not publish full user-agent strings. Search your logs for the token itself; under RFC 9309 a crawler's robots.txt token should appear as a substring of its user-agent header.
Does blocking ClaudeBot also block Claude-SearchBot?
No. RFC 9309 says a crawler must find the group that matches its own token and obey that group's rules. A User-agent: ClaudeBot group applies to ClaudeBot only. Claude-SearchBot and Claude-User fall back to your User-agent: * group unless they have groups of their own.
The mix-up runs in both directions. Some sites think they have blocked "Claude" with one ClaudeBot line and are surprised to see Claude-User in their logs. Others think a ClaudeBot block has removed them from Claude search, when it has only opted them out of training.
Cloudflare's managed robots.txt shows the intended pattern. Its generated file disallows ClaudeBot alongside GPTBot, Google-Extended, Applebot-Extended, CCBot and other training crawlers, and does not list Claude-SearchBot or Claude-User, with a content signal of search=yes, ai-train=no. ClaudeBot is also one of the most blocked agents on the web: Cloudflare's 2025 Radar review named GPTBot, ClaudeBot and CCBot as the AI crawlers most often fully disallowed in robots.txt.
What robots.txt groups should you use for each goal?
Pick the goal, then write one group per behaviour. Put the bots you allow in one group with your private-path rules repeated inside it, because a bot that matches a named group ignores the * group. The table maps each goal to its groups and what Claude can still do.
Your goal
ClaudeBot
Claude-SearchBot
Claude-User
Claude can still
Be cited, allow training
Allow
Allow
Allow
Train on, index and fetch your pages
Be cited, no training
Disallow
Allow
Allow
Index and fetch your pages
No training, no search index, allow user fetches
Disallow
Disallow
Allow
Read a page a user points it at
Out of Claude completely
Disallow
Disallow
Disallow
Mention you only from other sites or earlier training
The second row is the one most product sites want:
Tokens are matched case-insensitively but must be spelled exactly. claude-searchbot works; Claude-Search-Bot or ClaudeSearchBot matches nothing and leaves that crawler on your * rules. The robots.txt guide for AI bots has a full file with every operator in it.
What happens if you block all three?
Your own pages leave Claude's answers. Claude cannot index them for search, cannot fetch them when a user asks, and will not train on them from then on. Claude can still mention your product from what other websites say about you, if Claude-SearchBot indexed those pages, and from anything collected before you blocked it.
That is a legitimate choice for some sites, such as a publisher that licenses its content or a company with nothing to gain from AI referrals. For a product that wants to be recommended, it is expensive: every time a buyer asks Claude with web search on for tools in your category, your pricing page, docs and comparison pages are out of the pool Claude can quote. The guide on whether to block AI crawlers compares the costs across OpenAI, Anthropic, Perplexity, Google and Apple.
Do Anthropic's crawlers respect robots.txt and Crawl-delay?
Anthropic says yes: "Anthropic's Bots respect 'do not crawl' signals by honoring industry standard directives in robots.txt." It also supports Crawl-delay as a non-standard extension, so you can slow ClaudeBot down instead of blocking it. This differs from OpenAI and Perplexity, who say their user-triggered fetchers may skip robots.txt.
To throttle rather than block:
User-agent: ClaudeBot
Crawl-delay: 1
Three rules from Anthropic's help article save time:
Set robots.txt on every subdomain. Anthropic asks you to do this "for every subdomain that you wish to opt out from," so docs.example.com and app.example.com each need their own file.
Do not rely on IP blocking to opt out. Anthropic warns it "may not work correctly or persistently guarantee an opt-out, as doing so impedes our ability to read your robots.txt file."
If a crawler misbehaves, Anthropic lists a contact address in the same article for reports.
The contrast with the other operators is worth knowing when you write rules for all of them. OpenAI's documentation says of ChatGPT-User that "robots.txt rules may not apply," and Perplexity says Perplexity-User "generally ignores robots.txt rules" because a user requested the fetch. Anthropic draws no such exception for Claude-User.
How do Anthropic's crawlers line up with OpenAI's and Perplexity's?
The three operators split their crawlers the same way: one for training, one for search, one for fetches a user triggers. Knowing the equivalents lets you write one consistent policy instead of three. The table pairs them, with what each operator says about robots.txt for the user-triggered fetcher.
Behaviour
Anthropic
OpenAI
Perplexity
Training
ClaudeBot
GPTBot
None; Perplexity says PerplexityBot "is not used to crawl content for AI foundation models"
Search index
Claude-SearchBot
OAI-SearchBot
PerplexityBot
User-triggered fetch
Claude-User
ChatGPT-User
Perplexity-User
robots.txt for the user fetcher
Covered by Anthropic's statement that its bots honour robots.txt
"robots.txt rules may not apply"
"Generally ignores robots.txt rules"
A policy of "no training, yes citations" is therefore one Disallow group listing ClaudeBot, GPTBot and any other training crawlers you choose, and one allow group listing Claude-SearchBot, Claude-User, OAI-SearchBot, ChatGPT-User, PerplexityBot and Perplexity-User with your private paths repeated inside it. For the two user fetchers that may skip robots.txt, the only hard control is a firewall rule, and blocking them costs you the pages users paste into those assistants.
How do you verify Anthropic's crawlers in your logs?
Check the IP address against Anthropic's published list at claude.com/crawling/bots.json. The file has a creationTime and a prefixes array; in September 2026 it held 26 IPv4 prefixes, created in August 2026. It does not label which prefix belongs to which bot, so it tells you a request came from Anthropic, not which crawler sent it.
Pull the requests: grep -E "ClaudeBot|Claude-SearchBot|Claude-User" access.log.
Check the status codes each token received. A 200 means it read your page; 403, 429 or a challenge page means a firewall answered first.
Check a sample of source addresses against bots.json. A request claiming to be ClaudeBot from outside those prefixes is not Anthropic.
Reload bots.json on a schedule if you use it in firewall rules, since the prefixes change.
A robots.txt that looks right can still fail when a firewall, a CDN setting or a JavaScript-only page gets in the way. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. Its AI visibility checks flag a robots.txt that blocks Claude-SearchBot, OAI-SearchBot or PerplexityBot, fail a wildcard rule that blocks them by accident, treat a training-only block such as ClaudeBot as a legitimate pass, request your page as Claude-User to confirm it gets real content rather than a challenge or an empty shell, and check whether your main content is present in the raw HTML.
ClaudeBot collects web content that may be used to train Anthropic's models. Claude-SearchBot indexes content for Claude's search results. Blocking ClaudeBot keeps your future pages out of training; blocking Claude-SearchBot keeps them out of Claude's search.
03
How do I block ClaudeBot but still appear in Claude?
Give ClaudeBot its own robots.txt group with Disallow: / and give Claude-SearchBot and Claude-User a separate group that allows your public pages. Each bot follows only the group that names it.
04
Does Claude-User respect robots.txt?
Anthropic says its bots respect 'do not crawl' signals by honouring standard robots.txt directives, and it describes what disallowing Claude-User does. OpenAI and Perplexity, by contrast, say their user-triggered fetchers may not follow robots.txt.
05
Where does Anthropic publish its crawler IP addresses?
At claude.com/crawling/bots.json. In September 2026 the file listed 26 IPv4 prefixes and did not say which prefix belongs to which bot, so it verifies that a request came from Anthropic rather than which crawler sent it.
Perplexity cited content about 3x fresher than Google in one study. Put an accurate date in three places, visible, schema and sitemap, and never fake it.