Should you block AI crawlers? A decision table by bot
Block AI training crawlers if you want to; it costs no citations. Keep AI search bots and user agents open if you want AI answers to recommend you.
LLaunchScaler·Published ·8 min read
Block AI training crawlers if you do not want your content used to train models: it is a legitimate choice and it costs you no citations, because the operators run training and search as separate crawlers. Do not block the AI search crawlers or the agents that fetch pages for users if you want ChatGPT, Claude or Perplexity to recommend you, because blocking those removes your pages from their answers.
So "should I block AI crawlers?" has a different answer for each bot. The table below gives it bot by bot, using each operator's own description of what blocking does.
What are the three kinds of AI crawlers?
AI crawlers fall into three groups by what they do with your page. Training crawlers collect content for future models. Search crawlers index pages so an assistant can find and cite them. User-triggered agents fetch a page in real time because a person asked. Blocking each group costs you something different.
Future model training, and for Google-Extended, grounding in the Gemini app
Presence in what models learn; for Google-Extended, Gemini app grounding
Search
OAI-SearchBot, Claude-SearchBot, PerplexityBot
The index each assistant searches when answering
Citations in that assistant's answers
Questions, answered
What people ask about this
01
Should I block GPTBot?
Block it if you do not want your content used to train OpenAI's models; it costs you nothing in ChatGPT search, because OpenAI says its settings are independent. Keep OAI-SearchBot and ChatGPT-User allowed if you want ChatGPT to cite and read your pages.
02
Will blocking AI crawlers hurt my Google rankings?
Two of the training names are not crawlers at all. Google says Google-Extended has no user agent of its own and is only read "in a control capacity," and Apple says Applebot-Extended "does not crawl webpages." Both are switches that tell the operator how it may use pages its main crawler already fetched.
What does blocking each AI crawler cost you?
Each operator says in its own documentation what a block does. For training crawlers the cost is inclusion in future training; for search crawlers it is citations; for user agents it is the assistant's ability to read your page on request. The table quotes or summarises each operator and gives the usual choice for a product that wants to be recommended.
Token
Operator
Kind
What the operator says blocking does
For a product that wants recommendations
GPTBot
OpenAI
Training
Indicates your content "should not be used in training generative AI foundation models"
Your call; no effect on ChatGPT search
ClaudeBot
Anthropic
Training
Your "future materials should be excluded from our AI model training datasets"
Your call; no effect on Claude search
Google-Extended
Google
Training and Gemini app grounding
Opts out of Gemini training and grounding; "does not impact a site's inclusion in Google Search"
Your call; blocking also removes Gemini app grounding
Applebot-Extended
Apple
Training
Opts out of training Apple's foundation models; pages "can still be included in search results"
Your call
CCBot
Common Crawl
Crawl for an open dataset
Keeps your pages out of future crawls for Common Crawl's "free, open repository of web crawl data that can be used by anyone"
Your call
OAI-SearchBot
OpenAI
Search
Your site "will not be shown in ChatGPT search answers, though can still appear as navigational links"
Allow
Claude-SearchBot
Anthropic
Search
"May reduce your site's visibility and accuracy in user search results"
Allow
PerplexityBot
Perplexity
Search
Removes you from what it surfaces and links in Perplexity; Perplexity says it is not used for training
Allow
ChatGPT-User
OpenAI
User-triggered
OpenAI says "robots.txt rules may not apply" to it
Allow
Claude-User
Anthropic
User-triggered
"May reduce your site's visibility for user-directed web search"
Allow
Perplexity-User
Perplexity
User-triggered
Perplexity says it "generally ignores robots.txt rules"
Allow
Apart from Google-Extended's effect on the Gemini app, every "your call" row is a rights decision with no citation cost. Every "allow" row is a crawler that puts your pages in front of someone asking an assistant about your category.
Is blocking only training bots a legitimate setup?
Yes. The operators built their crawlers so you could do exactly this. OpenAI's documentation gives it as the example of independent settings: a site "can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot." Anthropic, Google and Apple describe their training controls as separate from search in the same way.
A training-only block keeps live citations intact because the assistants find pages to cite through their search crawlers and user agents, not through training crawls. It does have a cost, just not a citation one: with web search off, an assistant answers from what it learned in training, and a site that opts out is absent from that. For a young product that is rarely the deciding factor, since the answers that name and link products come from search. The comparison of OAI-SearchBot and GPTBot and the comparison of ClaudeBot and Claude-SearchBot show the exact groups for each operator.
Apple adds a wrinkle worth knowing. Its documentation says Applebot-crawled data "may be used to provide additional context and up-to-date content" when AI models generate answers in Siri and Search, answers that "may include links to sources and websites." Blocking Applebot-Extended does not opt you out of those; Apple says the control for them is the nosnippet meta tag on specific content. Even with both in place, Apple says your content "will remain discoverable through Spotlight, Siri, and Safari." So the Apple equivalent of "no training, yes citations" is a disallow for Applebot-Extended and no nosnippet on the pages you want quoted.
How much do training crawlers crawl compared with search crawlers?
Much more. Cloudflare's 2025 Radar review found that "crawling for model training is responsible for the overwhelming majority of AI crawler traffic, reaching as much as 7-8x search crawling and 32x user action crawling at peak." Blocking training crawlers therefore removes most AI crawl load from your server while leaving citations untouched.
Cloudflare's August 2025 analysis put numbers on the shares. Training drove "nearly 80% of AI bot activity, up from 72% a year ago," while search made up 18% and user actions 2% over the same twelve months. Between July 2024 and July 2025, search's share fell from 26% to 17%.
The traffic coming back is lopsided too. The same Radar review found Anthropic's crawl-to-refer ratio reached as much as 500,000 to 1 during 2025, OpenAI's as much as 3,700 to 1 in March, and Perplexity's generally stayed below 400 to 1. If server load from AI bots is your concern, the training crawlers are where it comes from.
When should you block AI search bots or user agents?
Block them only when you do not want an assistant to quote or read the content at all. Licensed or paywalled content, a publication whose business depends on people visiting, or legal limits on redistribution are the usual reasons. For a product site that wants buyers to find it, blocking them trades away the answers that could recommend you.
If the reason is a private area of the site, robots.txt is the wrong tool. RFC 9309 says robots.txt rules "are not a form of access authorization," and Cloudflare's documentation adds that compliance "is voluntary." Two operators say their user-triggered agents may not follow robots.txt at all. Anything that must stay private belongs behind a login.
What does a sensible robots.txt look like?
Write one group that disallows the training tokens you have chosen to block, and one group that allows the search crawlers and user agents with your private paths repeated inside it. A bot that matches a named group ignores your User-agent: * group, so rules you want everyone to follow must appear in each group.
Leave Google-Extended out of the first group if you want the Gemini app to keep grounding answers on your pages, since Google's token covers that as well as training. The robots.txt guide for AI bots explains each line and the mistakes that silently break it.
Can you block AI search crawlers without meaning to?
Yes, and the cause can sit outside robots.txt. A User-agent: * rule with Disallow: / catches every bot that lacks its own group. Firewall and CDN settings answer before robots.txt is read at all. Check both before assuming your policy is what you wrote.
Cloudflare is worth checking closely in 2026. Its documentation says that from September 15, 2026, new domains block bots classified as Training or Agent on pages that display ads, and that "mixed-purpose crawlers that combine Search and Training will also be blocked by all configurations to block AI training." Its managed robots.txt, by contrast, disallows training tokens such as GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot and leaves search crawlers alone. The guide to Cloudflare blocking AI crawlers walks through each setting.
Check what your site actually allows
What you meant to allow and what bots receive can differ, because of a wildcard rule, a CDN setting or a page that only renders with JavaScript. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. Its AI visibility checks fail a robots.txt that blocks OAI-SearchBot, Claude-SearchBot or PerplexityBot, fail a wildcard rule that catches them, pass a training-only block as a legitimate configuration, and test whether a firewall challenges retrieval bots or serves ChatGPT-User and Claude-User a challenge instead of your page.
Blocking AI training tokens and bots does not. Google states that Google-Extended does not affect inclusion or ranking in Google Search. Blocking Googlebot itself would, which is why a firewall-level AI block that also catches Googlebot needs care.
03
Which AI crawlers should a SaaS site allow?
The search crawlers (OAI-SearchBot, Claude-SearchBot and PerplexityBot) and the user-triggered agents (ChatGPT-User, Claude-User and Perplexity-User). Those are the ones that put your pages into AI answers and read the pages users send them.
04
Does robots.txt actually stop AI crawlers?
Only the ones that choose to follow it. Cloudflare's documentation says robots.txt compliance is voluntary, OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it. Enforcement needs a firewall rule.
05
Do AI training crawlers crawl more than AI search crawlers?
Far more. Cloudflare's 2025 Radar review found crawling for model training reached as much as 7 to 8 times search crawling at peak, and training made up nearly 80% of AI bot activity in the year to July 2025.
GA4 has an AI Assistant channel since May 13, 2026. How to find ChatGPT and Perplexity visits, and build a custom AI channel group with the regex to paste.
AI answers lean on reviews, listings and comparisons because buyer questions are comparative. What the citation studies show, and which pages to be on.