Claude learns about a product three ways, one Anthropic crawler each. Allow Claude-SearchBot and Claude-User, serve real HTML, and earn mentions.
LLaunchScaler·Published ·8 min read
Claude can mention your product in three ways: from what it learned in training, from Anthropic's search index, or from a page it fetches live while answering someone. Each of those has its own Anthropic crawler (ClaudeBot, Claude-SearchBot and Claude-User), so to be cited you allow the two that feed live answers, serve your content as real HTML, and earn mentions on other sites for the training side.
That split is the most useful thing to understand. Many sites block "Claude" in robots.txt to keep their content out of training and accidentally block the crawlers that would have cited them.
How does Claude find information about a website?
Claude draws on its trained knowledge and, when web search is on, on pages it retrieves for the question. Anthropic runs three separate bots, and each one feeds a different part of that. The table maps them using Anthropic's own descriptions and its own words on what blocking each one costs.
How Claude knows about you
Anthropic's crawler
What Anthropic says it does
What blocking it does, per Anthropic
Trained knowledge
ClaudeBot
Collects web content "that could potentially contribute to their training"
Your "future materials should be excluded from our AI model training datasets"
Search index
Claude-SearchBot
Analyses content "to enhance the relevance and accuracy of search responses"
Prevents indexing "which may reduce your site's visibility and accuracy in user search results"
Questions, answered
What people ask about this
01
Which Anthropic crawler do I need to allow for Claude to cite my site?
Claude-SearchBot and Claude-User. Anthropic says blocking Claude-SearchBot may reduce your visibility in Claude's search results, and blocking Claude-User prevents Claude from retrieving your pages when a user's question calls for them.
02
Does blocking ClaudeBot stop Claude from mentioning my product?
Accesses websites when individuals ask Claude questions
Prevents retrieving your content in response to a user query, "which may reduce your site's visibility for user-directed web search"
The first row decides what Claude says with web search off. The second and third decide whether Claude can find, read and link your pages with web search on, which is where citations with links come from.
Does blocking ClaudeBot stop Claude from citing you?
No. Blocking ClaudeBot only keeps your future pages out of Anthropic's training datasets. Claude-SearchBot and Claude-User are separate user agents, and under the robots.txt standard a crawler obeys the group that names it, so a ClaudeBot disallow does not touch them. Only a group naming them, or a * group when they have none of their own, applies.
A robots.txt that opts out of training and keeps both citation paths open looks like this:
# Anthropic: training
User-agent: ClaudeBot
Disallow: /
# Anthropic: search index and user-requested fetches
User-agent: Claude-SearchBot
User-agent: Claude-User
Allow: /
Disallow: /app/
Disallow: /account/
Three details trip people up. The private-path rules have to be repeated inside the Claude-SearchBot group, because a bot that matches a named group ignores your User-agent: * rules entirely. Anthropic asks you to set robots.txt "for every subdomain that you wish to opt out from," so docs.example.com needs its own file. And Anthropic warns that blocking by IP address "may not work correctly or persistently guarantee an opt-out," because it stops them reading your robots.txt; they publish their crawler addresses at claude.com/crawling/bots.json for verification instead.
How do you confirm Anthropic's crawlers can reach you?
Your server logs answer this faster than any tool. Search the access log for each user-agent token and check the status codes it received. A crawler that only ever gets 403 or a challenge page has never read your content, whatever robots.txt says, because a firewall answers before robots.txt is consulted.
Search the log for each token: grep -E "ClaudeBot|Claude-SearchBot|Claude-User" access.log.
Count the status codes per token. Mostly 200 is healthy. A run of 403, 429 or 503 points at bot protection or rate limiting.
Take a few of the IP addresses and check them against claude.com/crawling/bots.json. A request that claims to be ClaudeBot from an address outside that list is someone borrowing the name, and blocking it costs you nothing.
If no Anthropic token appears at all over several weeks, fetch your own robots.txt and read it as each bot would, group by group.
If the log shows challenges, the fix is in the firewall or CDN, not in robots.txt: add an allow or skip rule for the verified addresses, then watch the codes turn to 200. A firewall rule keyed on the user-agent string alone lets anyone who copies the string through, so match on the published addresses as well.
Can Claude read JavaScript-rendered pages?
Plan on no. Vercel's study of AI crawler traffic found that ClaudeBot fetches JavaScript files, 23.84% of its requests, but does not execute them, so content that only appears after JavaScript runs is not seen. Anthropic documents no rendering for Claude-SearchBot or Claude-User either, so treat the raw HTML as all any of them read.
A client-rendered app built with a default Vite and React setup serves <div id="root"></div> and a script tag. To an agent that does not run the script, that is an empty page, however complete it looks in a browser.
Test it in one command. Pick a sentence that appears on your page in a browser and search the raw HTML for it:
curl -s -A "Claude-User" https://yoursite.com/pricing | grep -c "per seat per month"
A count of 0 means the sentence is not in the HTML Claude receives. The fix is server-side rendering, static generation or prerendering the public routes, so the words arrive in the first response. Do AI crawlers render JavaScript? covers the fix per framework.
How does a young brand get into what Claude knows?
With web search off, Claude answers only from training data, and a product launched after that data was collected, or mentioned on few pages, is either missing or described vaguely. The way in is being written about by name on pages Anthropic's crawler reads: listings, reviews, comparisons, documentation and community threads that describe what you do.
Anthropic does not publish how it weights sources, so the practical approach is coverage and consistency. Work through these in order:
Write one plain sentence that says what the product is, who it is for and what it costs, and use it verbatim on your home page and every profile you control.
Get listed where your category is discussed: launch directories, category directories and review sites, each with that same description.
Publish documentation and a public changelog on your own domain, in server-rendered HTML, so the facts about your product exist in text.
Earn mentions in comparison articles and community threads by being useful there, never by planting posts.
Keep ClaudeBot allowed if you want this to count. A site that blocks training has chosen to be absent from this path, which is a legitimate choice with a known cost.
Mentions on other sites matter more than volume on your own, because a model that sees the same description of you from several independent sources has a consistent answer to give.
Why do Claude's answers differ with web search on and off?
With web search off, Claude answers from training data and gives no source links. With it on, Anthropic says the response includes "direct citations to sources" and "source links for further reading." The same question can therefore name different products, or none, depending on the toggle, so test both.
To switch it, Anthropic's help says to click the "+" button in the lower left corner of the chat window and select Web search. On Team and Enterprise plans an owner has to enable it in organization settings first.
Read the two results as two different diagnoses:
Web search
Claude names you?
What it tells you
What to work on
Off
No
Training data holds too little about you
Third-party mentions and listings, ClaudeBot allowed
Off
Yes, but wrong details
Sources disagree about what you do
One consistent description everywhere
On
No
Your pages were not found or not fetchable
Claude-SearchBot and Claude-User allowed, server-rendered HTML, direct answers
On
Yes, with a link
The citation path works
Keep it working after every deploy
How do you track whether Claude mentions you?
Keep 15 to 20 real buyer questions, run each once a month in a fresh chat with web search off and again with it on, and log whether Claude names you, links you and describes you correctly. Record the competitors it names instead; they show where the mentions you lack already exist.
Do not judge Claude by referral traffic alone. Cloudflare's 2025 Radar review found Anthropic had the highest crawl-to-refer ratio of the major AI platforms, reaching as much as 500,000:1, so Claude reads far more than it sends back as clicks. A mention in an answer can matter to a buyer who never clicks. The guide to measuring AI visibility covers how to score mentions and citations over time.
See whether Claude names you for your category's questions
A monthly log covers a few questions. LaunchScaler's full audit samples once, for $19, one time: it puts 10 buyer questions for your category to 7 answer engines, Claude among them with ChatGPT, Perplexity, Gemini, Copilot, Google AI Overviews and AI Mode, and reads all 70 answers to show which name you and which competitors they name instead. It unlocks on your report after the free scan, which checks the access side for Claude directly: whether robots.txt blocks Claude-SearchBot, whether a firewall challenges retrieval bots, whether the Claude-User agent gets a real page rather than a challenge or an empty shell, and whether your main content is in the raw HTML. Run the free scan first, then open the full audit.
No. ClaudeBot collects content for training, and blocking it only excludes your future pages from training datasets. Claude-SearchBot and Claude-User each follow their own robots.txt group, so they keep working unless you block them too.
03
Can Claude read a JavaScript-rendered website?
Treat it as no. Vercel's crawler research found ClaudeBot fetches JavaScript files but does not execute them, and Anthropic documents no rendering for its other agents. Put your content in the server-rendered HTML.
04
How do I turn web search on in Claude?
Anthropic's help says to click the plus button in the lower left of the chat window and choose Web search. With it on, Claude's answer includes direct citations and source links; with it off, Claude answers from what it learned in training.
05
Why doesn't Claude know about my product?
Without web search, Claude answers from training data, which may predate your launch or hold few mentions of you. With web search on, it depends on whether your pages can be fetched and whether they answer the question directly.
ChatGPT answers from live search or from what the model already knows. How to get into each: allow OAI-SearchBot, serve HTML, earn listings, then measure.