robots.txt changed without you knowing: how to monitor it
One robots.txt line can block Google or AI search bots. Changes come from deploys, plugins and CDNs, so diff the file on a schedule and test each bot.
LLaunchScaler·Published ·8 min read
Monitor robots.txt by fetching it on a schedule, keeping every copy, and alerting when its content or its HTTP status changes, then testing the new version against Googlebot and the AI search bots you want to reach. One line can stop Google or ChatGPT's search crawler reading your whole site, and the change can come from a deploy, a CMS plugin or a CDN setting that never touches your repository.
The file is small and public, which makes it cheap to watch. What makes changes easy to miss is that none of them alter how your pages look.
Why is a robots.txt change so high-impact?
robots.txt is read by every compliant crawler before it fetches anything else, and a single rule applies to every URL it matches. Disallow: / under User-agent: * stops all of them reading every page. A rule meant for one folder, or a new group for one bot, can quietly change what Google or an AI search engine is allowed to fetch across the site.
Three details in Google's specification make small edits bigger than they look:
A crawler obeys only one group. Google's crawlers pick "the group with the most specific user agent that matches the crawler's user agent," and ignore the * group if a specific one exists. Adding a User-agent: Googlebot group with a single rule means Googlebot stops following everything in your * group.
When rules conflict, "Google uses the least restrictive rule," and otherwise the most specific path wins, so an added Allow or Disallow can override a rule you thought was in charge.
Content past 500 KiB is ignored: "Content which is after the maximum file size is ignored."
A blocked page does not vanish from Google immediately. Google says a disallowed page "can still be indexed if linked to from other sites," but "the search result won't have a description," and Google cannot read or update the content.
Questions, answered
What people ask about this
01
How do I monitor robots.txt for changes?
Fetch https://yoursite.com/robots.txt on a schedule, save each copy, and alert when the content or the HTTP status differs from the last one. After any change, test the new file against Googlebot and the AI search bots you care about, such as OAI-SearchBot, PerplexityBot and Claude-SearchBot.
For AI search, the stakes are the same per bot. OpenAI's OAI-SearchBot is "used to surface websites in search results in ChatGPT's search features," Perplexity's PerplexityBot is "designed to surface and link websites in search results on Perplexity," and Anthropic's Claude-SearchBot "navigates the web to improve search result quality." A wildcard rule written to stop AI training can block these search bots too; robots.txt for AI bots shows a file that separates them.
Where do unexpected robots.txt changes come from?
They come from three places: your own deploys, your CMS and its plugins, and your CDN or host. Only the first shows up in code review. The other two can change the file served at /robots.txt while your repository copy stays the same, which is why you monitor the live URL rather than the file in git.
Source
How it changes the file
Example
A deploy
A generated robots route or a static file changes with the release
A robots route that serves Disallow: / unless an environment variable is set
The CMS itself
The file is generated dynamically rather than stored
WordPress up to 5.2 served Disallow: / from its generated robots.txt when "Discourage search engines" was ticked
CMS plugins
An SEO plugin with a robots.txt editor rewrites the file
A plugin update or settings import replaces your custom rules
CDN managed settings
The CDN adds rules to the response
Cloudflare's managed robots.txt will "prepend our managed robots.txt before your existing robots.txt, combining both into a single response"
Host or proxy errors
The file returns 4xx or 5xx instead of 200
A misrouted path that returns 404, or an origin outage that returns 503
Cloudflare's managed robots.txt adds Disallow rules for AI crawlers including GPTBot, ClaudeBot, Google-Extended, CCBot, Applebot-Extended, Amazonbot, Bytespider and meta-externalagent. Cloudflare also has a separate bot-blocking setting that blocks requests at its edge whatever robots.txt says; its documentation says that setting targets training bots and "excludes mixed-purpose bots that are used both for Training and for Search." Both are switches someone on your team can flip in a dashboard. The firewall side is covered in Cloudflare blocking AI crawlers.
The status code matters as much as the content. Google treats 4xx errors other than 429 "as if a valid robots.txt file didn't exist," so a 404 opens everything. For a 5xx, Google pauses crawling for the first 12 hours, then uses its last cached copy for up to 30 days, and after that treats the site as having no robots.txt if the site is otherwise available.
How do you diff robots.txt on a schedule?
Fetch the live file on a timer, compare it with the previous copy, and alert on any difference in the body or the status code. Hourly is plenty for most sites, since Google generally caches the file "for up to 24 hours." Keep every version with a timestamp so you can line a change up with a traffic dip later.
A cron job and a short script are enough:
#!/usr/bin/env bash
# robots-watch.sh: run hourly from cron, e.g. 0 * * * * /path/robots-watch.sh
URL="https://yoursite.com/robots.txt"
DIR="$HOME/robots-watch"; mkdir -p "$DIR"
code=$(curl -s -o "$DIR/new.txt" -w '%{http_code}' "$URL")
if [ "$code" != "200" ]; then
echo "robots.txt returned $code" | mail -s "robots.txt status $code" you@yoursite.com
fi
if [ -f "$DIR/last.txt" ] && ! diff -q "$DIR/last.txt" "$DIR/new.txt" >/dev/null; then
diff "$DIR/last.txt" "$DIR/new.txt" | mail -s "robots.txt changed" you@yoursite.com
cp "$DIR/new.txt" "$DIR/robots-$(date +%Y%m%d-%H%M).txt"
fi
mv "$DIR/new.txt" "$DIR/last.txt"
Run it for every host that serves pages, because each host has its own file: Google's rules "apply only to the host, protocol, and port number where the robots.txt file is hosted." www.yoursite.com, yoursite.com, docs.yoursite.com and blog.yoursite.com each need a check.
Search Console gives you a second view for Google. Its robots.txt report shows "which robots.txt files Google found for the top 20 hosts on your site, the last time they were crawled, and any warnings or errors encountered." After fixing a bad file you can select Request a recrawl there, though Google notes that "you generally don't need to request a recrawl of a robots.txt file, because Google recrawls your robots.txt files often."
Which robots.txt changes should trigger an alert?
Alert on every change to the body, but rank them, so the dangerous ones are read first. A comment edit and a new Disallow: / look the same to a plain diff. Add a few pattern checks to the script so its subject line tells you which kind of change arrived.
Change
Why it matters
Priority
Status is not 200
A 4xx opens everything to Google; a 5xx pauses Google's crawling
Immediate
A Disallow: / line appears
One group now blocks the whole site
Immediate
A new User-agent: group appears
That bot now ignores your * group entirely
Same day
A retrieval bot (OAI-SearchBot, PerplexityBot, Claude-SearchBot) is named in a Disallow group
AI search stops reading those paths
Same day
The Sitemap: line disappears or changes
Crawlers lose the pointer to your sitemap
This week
The file differs between www and the apex host
Each host is governed by its own file
This week
Comments or whitespace only
No effect on crawling
Log it
How do you test Googlebot and AI bots against a new robots.txt?
After any change, check whether each bot you care about may fetch your important URLs, using a parser that follows Google's rules. Reading the file by eye is how most mistakes get through, because group selection and rule precedence are not obvious. Test a list of bots against a list of URLs and compare the answers with what you intended.
Google publishes the parser it uses in production as an open-source C++ library. Build it once, then run it against the live file:
git clone https://github.com/google/robotstxt.git && cd robotstxt
bazel build :robots_main
curl -s https://yoursite.com/robots.txt -o /tmp/robots.txt
for bot in Googlebot OAI-SearchBot ChatGPT-User PerplexityBot Claude-SearchBot GPTBot ClaudeBot; do
for url in https://yoursite.com/ https://yoursite.com/pricing https://yoursite.com/blog/; do
bazel-bin/robots_main /tmp/robots.txt "$bot" "$url"
done
done
Each line prints ALLOWED or DISALLOWED, and the exit code is 0 for allowed and 1 for disallowed, so the loop can fail a pipeline. The results you usually want for a SaaS site:
Bot
Operator's stated purpose
Usual intent for a public marketing site
Googlebot
Google Search
Allowed
OAI-SearchBot
Surfacing sites in ChatGPT search
Allowed
ChatGPT-User
User-initiated actions in ChatGPT
Allowed; OpenAI says robots.txt rules "may not apply"
PerplexityBot
Surfacing and linking sites in Perplexity
Allowed
Claude-SearchBot
Improving search result quality for Claude users
Allowed
GPTBot
Crawling content for model training
Your choice
ClaudeBot
Collecting content for model training
Your choice
Robots rules only matter if the bot reaches the file. A firewall that challenges a bot blocks it regardless of what robots.txt says, so test the page response too: curl -s -o /dev/null -w '%{http_code}' -A "OAI-SearchBot" https://yoursite.com/. A 403 or a challenge page means the block is in your firewall, not your file. For a starting file with sensible defaults, see the robots.txt example for a SaaS site.
What should you do when robots.txt changed unexpectedly?
Find the source before you edit the file, or the next deploy, plugin sync or CDN setting will put the change back. Then restore the intended rules, confirm with the parser, and let the crawlers pick the file up. Google, OpenAI and Perplexity all describe roughly a day for a new file to take effect.
Compare the new file with the last good copy to see exactly which lines changed.
Check the repository history for the robots file or route, the CMS plugin settings and the CDN dashboard, in that order.
Restore the intended rules at the source, then fetch the live file to confirm it is served with a 200.
Re-run the bot and URL tests.
In Search Console's robots.txt report, request a recrawl if the bad version blocked important pages.
Record the incident with its time window, so a later dip in Google or AI referrals can be matched to it.
Watch robots.txt and the rest of the site together
An hourly diff catches the file changing. It does not tell you whether the pages it governs still send a noindex, whether your firewall now challenges AI search bots, or whether the site is up. LaunchScaler Watch re-scans one website every two weeks for $29/mo, including the scan's checks for robots.txt blocking pages you want ranked and for AI retrieval bots blocked by a rule or a firewall, emails a diff on every run with the evidence behind each change, runs uptime checks on your health endpoint in between, and alerts you when a check drops below the bar. Start watching.
How quickly do search engines pick up a robots.txt change?
Google generally caches robots.txt for up to 24 hours. OpenAI says it can take about 24 hours for its systems to adjust after an update, and Perplexity says it may take up to 24 hours.
03
What can change my robots.txt without a code deploy?
CMS plugins with a robots.txt editor, a CMS that generates the file dynamically, and CDN settings. Cloudflare's managed robots.txt, for example, prepends its own rules for AI crawlers before your existing file and serves both as one response.
04
What happens if robots.txt returns an error?
Google treats a 4xx response other than 429 as if no robots.txt existed. For a 5xx it stops crawling for the first 12 hours, uses the last cached copy for up to 30 days, and after that treats the site as having no robots.txt if the site is otherwise available.
05
Can one line in robots.txt remove my site from Google?
A Disallow: / under User-agent: * stops every compliant crawler reading your pages. Google may keep URLs it already knows in results without a description, but it can no longer read the content, and new pages cannot be crawled.
SEO monitoring for a small site: uptime every few minutes, indexing directives on every deploy, a technical scan every two weeks, Search Console monthly.