robots.txt for AI bots: block training, keep citations
A copy-paste robots.txt that disallows AI training crawlers and allows AI search and user agents, one commented group per bot, plus the traps to avoid.
LLaunchScaler·Published ·9 min read
A robots.txt that blocks AI training and keeps AI citations gives each training token (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot) its own group with Disallow: /, and gives each search and user agent (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User and PerplexityBot) a group that allows your public pages. The file below does that with one commented group per bot, ready to copy.
Two traps can undo it without any visible error: a wildcard "block all" rule that catches the search crawlers along with everything else, and a misspelled token that silently matches nothing.
What should robots.txt say to AI bots?
Disallow the training crawlers you do not want, allow the crawlers that feed AI answers, and repeat your private-path rules inside every allow group. A bot that finds a group naming it ignores the User-agent: * group, so the paths you hide from everyone must be listed again for each named bot.
Replace /app/, /account/ and /api/ with your own private paths, and the sitemap URL with yours.
# ---- AI training: opt out ----
# OpenAI: collects content that may be used to train its models
User-agent: GPTBot
Disallow: /
# Anthropic: collects content that may be used for training
User-agent: ClaudeBot
Disallow: /
# Google: token only; covers Gemini training and Gemini app grounding.
# Not Google Search, AI Overviews or AI Mode. Remove this group if you
# want the Gemini app to ground answers on your pages.
User-agent: Google-Extended
Disallow: /
# Apple: token only; opts out of training Apple's foundation models
User-agent: Applebot-Extended
Disallow: /
# Common Crawl: builds an open dataset anyone can use
User-agent: CCBot
Disallow: /
# ---- AI search and user fetches: allow ----
# OpenAI: ChatGPT search results
User-agent: OAI-SearchBot
Allow: /
Disallow: /app/
Disallow: /account/
Disallow: /api/
# OpenAI: pages fetched when a ChatGPT user asks
User-agent: ChatGPT-User
Allow: /
Disallow: /app/
Disallow: /account/
Disallow: /api/
# Anthropic: Claude's search index
User-agent: Claude-SearchBot
Allow: /
Disallow: /app/
Disallow: /account/
Disallow: /api/
# Anthropic: pages fetched when a Claude user asks
User-agent: Claude-User
Allow: /
Disallow: /app/
Disallow: /account/
Disallow: /api/
# Perplexity: surfaces and links sites in Perplexity search
User-agent: PerplexityBot
Allow: /
Disallow: /app/
Disallow: /account/
Disallow: /api/
# ---- Everyone else, including Googlebot and Bingbot ----
User-agent: *
Allow: /
Disallow: /app/
Disallow: /account/
Disallow: /api/
Sitemap: https://www.example.com/sitemap.xml
Questions, answered
What people ask about this
01
How do I block AI training bots in robots.txt without losing AI search?
Give each training token (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot) a group with Disallow: /, and give each search and user agent (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot) a group with Allow: / and your private paths disallowed.
Each group's purpose comes from the operator's own documentation. OpenAI says disallowing GPTBot indicates content "should not be used in training generative AI foundation models," and that sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers." Anthropic says blocking ClaudeBot excludes "future materials" from training datasets. Apple says Applebot-Extended "does not crawl webpages" and only governs how Applebot's data is used. Google says Google-Extended "does not impact a site's inclusion in Google Search."
Perplexity-User is not in the file on purpose. Perplexity says it "generally ignores robots.txt rules" because a user requested the fetch, so a group for it changes little; it falls under * either way. The guide on whether to block AI crawlers covers when you might choose differently for each bot.
How does a crawler decide which group applies?
A crawler looks for the group whose User-agent line matches its own token, matching case-insensitively, and obeys only that group. If several groups match, their rules are merged. If none match, it obeys the User-agent: * group. Within the group, the longest matching path wins. Those four rules, from RFC 9309, explain every line above.
Rule, per RFC 9309
What it means for your file
Crawlers "MUST use case-insensitive matching to find the group that matches the product token"
claudebot and ClaudeBot both work; spelling must be exact
Several groups matching the same token "MUST be combined into one group"
A second GPTBot group further down adds to the first rather than replacing it
With no matching group, crawlers "MUST obey the group with a user-agent line with the '*' value"
Bots you do not name follow your * rules
"The most specific match found MUST be used," measured in characters
Disallow: /app/ beats Allow: / for URLs under /app/; on a tie, Allow should win
The first and third rules together are why private paths are repeated. The OAI-SearchBot group above is the only one OAI-SearchBot reads, so without its own Disallow: /app/ it would be allowed into /app/.
A group can also carry several User-agent lines above one set of rules, which shortens the file. One group per bot, as above, is longer but easier to audit, and the comments tell the next person why each rule exists.
Why does a "block all AI" rule also block citations?
Because robots.txt has no category for "AI." A rule meant to keep AI out either lists every bot by name, and can end up listing the search crawlers too, or uses the * wildcard, which applies to every crawler that lacks its own group, AI search crawlers and search engines included. Either way, the bots that produce citations get blocked along with the ones that train.
What the file does
What happens
Fix
User-agent: * then Disallow: /, with a Googlebot group that allows
OAI-SearchBot, Claude-SearchBot and PerplexityBot have no group, fall back to * and are blocked
Add allow groups for the search and user agents
A "block AI" list that includes OAI-SearchBot or PerplexityBot
Those engines stop citing you
Keep the list to training tokens
User-agent: *bot* or User-agent: AI*
RFC 9309 tokens contain only letters, underscores and hyphens, or a lone *; these match no crawler
Name each bot
A block pasted from a list that names ClaudeBot only, meant to cover Claude
Claude-SearchBot and Claude-User are untouched, or caught by * if it disallows
Decide per bot and write each group
LaunchScaler's scan treats this as a failed check rated critical: it fails any "block all AI" or * rule that also matches OAI-SearchBot, Claude-SearchBot or PerplexityBot without an allow group carving them out, because it silently removes a site from AI answers.
What if you really want to disallow all AI bots?
Then name every one of them, search crawlers and user agents included, and accept what it costs: no citations in ChatGPT search, Claude or Perplexity, and no Gemini app grounding. Do not reach for User-agent: *, because that also blocks Googlebot and Bingbot and takes you out of search results entirely.
This still leaves you in Google Search and its AI features, because Google-Extended does not govern them; Google's snippet controls do. It also leaves the two user-triggered fetchers that OpenAI and Perplexity say may skip robots.txt, so if those must be kept out, the block belongs in your firewall, matched against each operator's published IP list. New AI crawlers appear regularly, and a named list only covers the names on it, so a policy like this needs rechecking against each operator's documentation.
Should you add Crawl-delay?
Only if a crawler's request rate is a real problem for your server. Crawl-delay is not part of RFC 9309, and support varies by operator. Anthropic says it supports "the non-standard Crawl-delay extension to robots.txt," with Crawl-delay: 1 under User-agent: ClaudeBot as its example. Slowing a training crawler this way is gentler than blocking it, and it leaves the search crawlers untouched.
What happens if you misspell a bot's token?
The group matches nothing, and the bot falls back to your * group as if you had written no rule for it. There is no error and no warning. Case does not matter, but every letter and hyphen does, so copy tokens from the operator's own documentation rather than retyping them.
No. robots.txt is a request, not a lock. RFC 9309 says its rules "are not a form of access authorization," and Cloudflare's documentation puts it plainly: "robots.txt compliance is voluntary." The operators named here say they follow it, except for user-triggered fetchers, which OpenAI and Perplexity say may not.
The opposite problem matters more for citations. A firewall answers a request before robots.txt is consulted, so a WAF rule or bot-protection setting that challenges OAI-SearchBot blocks it however open your file is. Cloudflare's AI Crawl Control, for example, enforces its blocks with a WAF custom rule, and its documentation warns that allowed crawlers "may still be affected by other security rules that execute before" that rule. Read the guide to Cloudflare blocking AI crawlers if your site sits behind Cloudflare.
How do you deploy and test the file?
Serve it as plain text at /robots.txt on every host, confirm it returns 200, then read the live copy rather than the one in your repository. A CDN, framework or plugin can serve something different from what you committed, and a server error on this one URL can block every crawler from the whole site.
Put the file at the root of each host: www.example.com/robots.txt, docs.example.com/robots.txt and so on. Anthropic asks for this "for every subdomain that you wish to opt out from."
Check the status: curl -s -o /dev/null -w "%{http_code}\n" https://www.example.com/robots.txt. It must be 200. RFC 9309 says a 500-range error means crawlers "MUST assume complete disallow."
Read what is served: curl -s https://www.example.com/robots.txt. Look for lines you did not write, such as a # BEGIN Cloudflare Managed content block that Cloudflare prepends when its managed robots.txt is on.
Keep the file well under 500 KiB, the minimum RFC 9309 requires crawlers to parse.
Allow time. OpenAI says robots.txt changes can take about 24 hours to reach its systems, and Perplexity says up to 24 hours.
For a full SaaS file that also handles staging, search pages and parameters, see the robots.txt example for SaaS.
Check the file against every AI crawler at once
Reading robots.txt as ten different bots is slow, and a firewall rule is invisible in the file. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. It reads your live robots.txt for OAI-SearchBot, Claude-SearchBot and PerplexityBot, passes a training-only block, requests your pages with retrieval bots' and agents' user agents to catch a 403 or a challenge, and flags a robots.txt rule that blocks a page you want ranked in Google.
Can one robots.txt line block every AI crawler?
No. robots.txt has no wildcard for 'AI bots'. The only wildcard is User-agent: *, which applies to every crawler without its own group, Googlebot and the AI search crawlers included, so Disallow: / under it blocks search engines too.
03
Is the user-agent in robots.txt case-sensitive?
No. RFC 9309 says crawlers must match the product token case-insensitively, so claudebot and ClaudeBot both work. The spelling must still be exact: ClaudeSearchBot or OAI-Search-Bot matches nothing.
04
Do AI crawlers obey robots.txt?
The training and search crawlers from OpenAI, Anthropic and Perplexity say they do. OpenAI says robots.txt may not apply to ChatGPT-User and Perplexity says Perplexity-User generally ignores it, because a user triggered the fetch.
05
What happens if my robots.txt returns a server error?
RFC 9309 says a crawler that gets a 500-range error for robots.txt must assume complete disallow. A robots.txt that errors for bots, for example behind a firewall, can block every crawler from the whole site.
Google says AI features need no special schema. What controlled studies show about schema and AI citations, and the six types a SaaS site actually needs.
GA4 has an AI Assistant channel since May 13, 2026. How to find ChatGPT and Perplexity visits, and build a custom AI channel group with the regex to paste.