Cloudflare blocking AI crawlers: check it and allow the right ones
Cloudflare can block AI crawlers before robots.txt is read. Where each setting lives, how to test with curl, and how to allow AI search bots only.
LLaunchScaler·Published ·8 min read
Cloudflare can block AI crawlers before your robots.txt is ever read, so a site whose robots.txt allows ChatGPT, Claude and Perplexity can still be invisible to them. Since July 2025 Cloudflare has asked new domains at signup whether to allow AI crawlers, and from September 15, 2026 new domains block Training and Agent bots on pages that show ads by default. The fix is to find which setting is answering the crawlers, then allow the search crawlers and user agents while blocking only training.
Older zones are not exempt. If yours already blocked AI training, through the newer options or the legacy "Block AI bots" toggle, Cloudflare now applies that block to crawlers that do both search and training, Googlebot included, unless someone opted out before September 15, 2026. The steps below cover every Cloudflare switch involved, checked against Cloudflare's documentation in September 2026.
Does Cloudflare block AI crawlers by default?
For new domains, partly, and the default has changed twice. On July 1, 2025 Cloudflare began asking every new domain at signup whether it wants to allow AI crawlers, starting it with "the default of control." On September 15, 2026, new domains began blocking Training and Agent bots on ad-bearing pages by default, while Search stays allowed.
Date
What changed
Who it affects
July 1, 2025
New domains are asked at signup whether to allow AI crawlers
New domains
July 1, 2026
AI bot policies split AI traffic into Search, Agent and Training, each with its own setting, for all customers including Free
Every zone, when you choose to use them
Questions, answered
What people ask about this
01
Does Cloudflare block AI crawlers by default?
For new domains, partly. Since July 2025 Cloudflare asks every new domain at signup whether to allow AI crawlers. From September 15, 2026, new domains block bots classified as Training or Agent on pages that display ads by default, while Search stays allowed.
New defaults: Training and Agent blocked on pages that display ads, Search allowed. Any Training block, legacy toggle included, now also blocks crawlers that combine search and training. The legacy "Block AI bots" toggle is deprecated
New domains for the defaults; every zone that blocks Training for the mixed-purpose rule, unless it opted out beforehand
The rule that reaches existing zones is the mixed-purpose one. Cloudflare's documentation states it: "Mixed-purpose crawlers that combine Search and Training will also be blocked by all configurations to block AI training, including the legacy 'Block AI bots' option." Its July 2026 announcement names the crawlers it means: "multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service)." Owners could opt out in Security settings before September 15 to keep Training crawlers that also crawl for Search unchanged. If you are not sure whether anyone did, check Security Settings now and look for Googlebot in your Security Events.
Which Cloudflare settings block AI crawlers?
Six places can stop an AI crawler, and most sit under Security Settings in the dashboard. Each one acts differently: some block at the firewall, one only writes robots.txt rules, and one issues challenges that cannot be skipped. Check all six before assuming Cloudflare is not involved.
Setting
Where it is
What it does to AI crawlers
AI bot policies
Security Settings > Configure AI bot policies
Search, Agent and Training each set to "Block (on all pages)", "Block on pages with ads" or "Allow (do not block)". Blocks verified bots with that behaviour plus similar unverified bots
Block AI bots (legacy)
Security Settings > Block AI bots
Blocks verified AI training bots and similar unverified bots; deprecated September 15, 2026
Managed robots.txt
Security Settings, filter by Bot traffic, "Set your preference to block training in robots.txt"
Prepends Disallow groups for Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent, plus Content-signal: search=yes, ai-train=no, use=reference
AI Crawl Control
AI Crawl Control in the dashboard
Allow or Block per named crawler; blocks are enforced by one WAF custom rule named "AI Crawl Control"
Bot Fight Mode
Security Settings, filter by Bot traffic, "Bot fight mode"
Challenges traffic matching known bot patterns; cannot be bypassed with WAF custom rules
WAF custom rules
Security rules, filtered by Custom rules
Any rule you or a teammate wrote that blocks or challenges bots, countries or user agents
The managed robots.txt is the gentle one: it only asks, and it leaves search crawlers alone. The AI bot policies, AI Crawl Control and the WAF enforce, and Bot Fight Mode enforces in a separate pipeline where Cloudflare says "Skip, Bypass, and Allow actions have no effect."
Why can Cloudflare block AI bots that robots.txt allows?
Because Cloudflare answers the request before it reaches your site, and robots.txt only matters to a crawler that got through. A WAF rule, an AI bot policy or a bot challenge returns a 403 or a challenge page, and the crawler never sees your content. Cloudflare's own robots.txt documentation calls robots.txt compliance "voluntary"; a firewall block is not.
Cloudflare's docs spell out the order: WAF custom rules, including AI Crawl Control's blocks, run first, then Cloudflare's bot solutions. The AI Crawl Control rule is added at the end of your existing custom rules, so an earlier rule that blocks bots wins over an Allow you set later. Cloudflare's troubleshooting note says allowed crawlers "may still be affected by other security rules that execute before the AI Crawl Control rule."
A challenge is as final as a block for these crawlers. Anthropic says its bots "respect anti-circumvention technologies" and "will not attempt to bypass CAPTCHAs." OpenAI recommends "allowing requests from our published IP ranges" for OAI-SearchBot, which is the firewall half of being citable. The guide to fixing "Blocked due to access forbidden (403)" covers the same problem as Googlebot sees it.
How do you test whether Cloudflare blocks ChatGPT, Claude or Perplexity?
Request a page with each AI crawler's user agent and read the status code and headers. A 200 with your content is what you want. A 403, or any response carrying cf-mitigated: challenge, means Cloudflare answered instead of your site. Then confirm in Cloudflare's event log, because a curl from your laptop is not the real bot.
Send a request as OAI-SearchBot and print the status and the challenge header:
curl -s -o /dev/null -D - -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://yoursite.com/ | grep -iE "^HTTP|cf-mitigated"
Repeat with PerplexityBot/1.0, ChatGPT-User/1.0, Claude-User and Claude-SearchBot in the user-agent string.
Read the result. Cloudflare says a challenge page always carries cf-mitigated: challenge and is served as text/html whatever you requested.
In the dashboard, open Security > Analytics, select the Events tab and filter for the crawler. The Service field names the feature that acted, and Bot Fight Mode's actions are labelled "Bot Fight Mode."
Open AI Crawl Control and read the Crawlers tab: each AI crawler's allowed and unsuccessful requests, and its robots.txt violations.
Read step 1 with care. Your request comes from your IP, not the operator's, so a rule keyed on verified bots treats it differently from the real crawler, and on the Free plan AI Crawl Control identifies crawlers by user-agent string alone. The Security Events log, filtered to the operator's published addresses, is the authoritative answer. The guide to checking whether AI bots crawl your site covers reading those logs.
How do you allow AI search bots and still block training?
Keep enforcement for the search crawlers and user agents on Allow, and express the training opt-out in robots.txt, or by blocking the named training crawlers individually. Avoid a firewall-wide Training block: Cloudflare says it also catches mixed-purpose crawlers, and its announcement lists Googlebot among them.
In Security Settings > Configure AI bot policies, set Search to "Allow (do not block)" and Agent to "Allow (do not block)." Set Training to "Allow (do not block)" too, unless you have confirmed your search engine traffic is unaffected by blocking it.
Turn on the managed robots.txt ("Set your preference to block training in robots.txt"), or publish your own groups that disallow GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot and allow OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User and PerplexityBot.
If you want training blocked at the firewall, open AI Crawl Control and set the individual training crawlers you have chosen to Block, leaving the search crawlers and user agents on Allow.
Review Security rules, filtered by Custom rules. Move or edit any rule that blocks or challenges bots so it does not catch the crawlers you allow. Cloudflare's example for sparing known good bots uses the Known Bots field, cf.client.bot, and a rule with the Skip action lets matched requests bypass other security features.
If Security Events show Bot Fight Mode acting on requests from an AI search crawler's published addresses, note that it "cannot be customized, adjusted, or reconfigured via WAF custom rules." Cloudflare's documented path for exceptions is Super Bot Fight Mode, which supports Skip rules; otherwise turn Bot Fight Mode off.
Re-run the curl tests and check Security Events for the crawlers a day later. OpenAI says robots.txt changes can take about 24 hours to reach its systems.
Cloudflare's dashboard shows what it did; it does not show what each AI engine ended up able to read. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. Its AI visibility checks request your pages with the retrieval bots' user agents and fail a 403 or a challenge that a browser never sees, request them as ChatGPT-User and Claude-User to confirm an agent gets a real page, read robots.txt for OAI-SearchBot, Claude-SearchBot and PerplexityBot, and pass a training-only block as a legitimate choice.
In the Cloudflare dashboard, go to Security Settings and choose Configure AI bot policies, where Search, Agent and Training each have Block, Block on pages with ads, or Allow. The older Block AI bots toggle, in the same place, is being deprecated on September 15, 2026.
03
Will blocking AI training in Cloudflare also block Googlebot?
It can. Cloudflare's documentation says mixed-purpose crawlers that combine Search and Training are blocked by all configurations that block AI training, and its July 2026 announcement names Googlebot, Applebot and BingBot as such crawlers.
04
How do I allow OAI-SearchBot and PerplexityBot through Cloudflare?
Set the Search behaviour to Allow in AI bot policies, set those crawlers to Allow in AI Crawl Control, and check that no WAF custom rule or Bot Fight Mode challenges them. Then request a page with each crawler's user agent and look for a 200.
05
Does Cloudflare's managed robots.txt block ChatGPT search?
No. The managed robots.txt disallows training crawlers such as GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot, and sets the content signal search=yes. It does not add a group for OAI-SearchBot, Claude-SearchBot or PerplexityBot.
Perplexity cited content about 3x fresher than Google in one study. Put an accurate date in three places, visible, schema and sitemap, and never fake it.