Blocked due to access forbidden (403): unblocking Googlebot
A 403 to Googlebot usually comes from bot protection, not your app. Find the layer in your WAF logs and allow verified crawlers by IP, never by user agent.
LLaunchScaler·Published ·8 min read
Blocked due to access forbidden (403) means your server answered Googlebot with HTTP 403 Forbidden, so Google could not crawl the page. On pages you want indexed, the 403 usually comes from bot protection in front of your app, such as a CDN bot rule, a WAF, a rate limit or geo-blocking, and the fix is to let verified crawlers through without opening the site to scrapers.
Most of the work is finding which layer answered. Your app logs often show nothing, because the request never reached the app.
What does "Blocked due to access forbidden (403)" mean?
It means the server refused Googlebot. Search Console's help says: "HTTP 403 means that the user agent provided credentials, but was not granted access. However, Googlebot never provides credentials, so your server is returning this error incorrectly. The page will not be indexed." Its advice is to admit non-signed-in users, or "explicitly allow Googlebot requests without authentication."
Google's crawler documentation adds that it "doesn't use the content from URLs that return 4xx status codes," and that 404 and 410 cause a previously indexed URL to be removed from the index. It also says not to use 403 as a throttle: "Don't use 401 and 403 status codes for limiting the crawl rate. The 4xx status codes, except 429, have no effect on crawl rate."
A 403 on a page you meant to keep private is fine. Open the row, read the example URLs, and act only on the ones that should be in search. The page indexing report guide covers how this row relates to the others.
Which layer is returning the 403?
Walk the request path from the outside in: CDN or bot protection, then the WAF, then rate limiting, then geo rules, then the app's own auth. The first layer that logs a block for the failing requests is the one. The response body often gives it away too, because a CDN challenge page looks nothing like your app's error page.
Layer
Typical trigger
Where to look
Questions, answered
What people ask about this
01
What does blocked due to access forbidden (403) mean?
Your server answered Googlebot with HTTP 403 Forbidden, so Google could not crawl or index the page. Google's help notes that Googlebot never provides credentials, so a 403 to Googlebot means the server is returning it incorrectly for a page you want indexed.
02
How do I fix blocked due to access forbidden (403)?
A bot-fighting mode or "block AI bots" setting that challenges automated traffic
The CDN's security event log, filtered by the service that acted
WAF rule
A managed or custom rule matching the crawler's user agent, path or request rate
WAF logs for the failing URLs and times
Rate limiting
Crawl bursts exceeding a per-IP limit
Rate-limit logs; the response may be 403 or 429
Geo-blocking
A country block that includes where the crawler connects from
Firewall country rules
App or host auth
A password-protected deployment or a login middleware
App logs, host project settings
Two quick checks narrow it down before you open any dashboard:
In URL Inspection, run Test live URL on an affected page. Open View tested page and read the HTTP response and the HTML. A challenge page or a CDN-branded error means the block is at the edge.
Compare the status your browser gets with the status a plain request gets:
If curl gets 403 and your browser gets 200, the rule is reacting to non-browser traffic in general, which is exactly what crawlers are.
How do you find the Googlebot requests in Cloudflare or your WAF?
Search the security log for the failing URLs around their crawl times, then read which rule or service acted. In Cloudflare, Security Events lists requests its security products blocked or challenged, with top sources by IP, user agent, path and country, plus sampled logs showing the action and the product that took it.
Get the affected URLs and their last crawl dates from the Page indexing report or from Crawl Stats.
Open your CDN's or WAF's security event log for that window. In Cloudflare that is Security Events; the Free plan shows sampled logs only, while Pro and above show the full dashboard.
Filter by path, or by user agent containing Googlebot.
Read the service and rule that acted. Cloudflare labels requests challenged by Bot Fight Mode with "Bot Fight Mode" in the Service field.
Check the source IPs of the blocked requests against Google's ranges before changing anything, since requests with a Googlebot user agent are not always Google.
One Cloudflare detail changes the fix. Its docs say "you cannot bypass or skip Bot Fight Mode using WAF custom rules or Page Rules," because it runs in a separate pipeline. If Bot Fight Mode is the layer blocking crawlers you need, turn it off, or move to Super Bot Fight Mode, which "runs on the Ruleset Engine and supports Skip rules."
How do you allow verified crawlers without opening the site?
Allow them by identity you can verify, never by user agent string. Anyone can send Googlebot/2.1 in a header, so a rule that trusts the user agent lets every scraper that copies it walk straight past your protection. Google gives two ways to verify its crawlers, and most CDNs expose their own verified-bot flag.
Google's manual check is a reverse DNS lookup on the IP, confirming the hostname ends in googlebot.com, google.com or googleusercontent.com, then a forward lookup confirming the name maps back to the same IP:
host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1
For firewall rules, match against the IP ranges Google publishes in CIDR format. Googlebot is in common-crawlers.json at https://developers.google.com/static/crawling/ipranges/common-crawlers.json. The older googlebot.json path still exists but now redirects there (checked in September 2026), so a script that downloads it without following redirects gets an HTML redirect page instead of the ranges.
On Cloudflare, the cf.client.bot field "indicates whether the request originated from a known good bot or crawler," and cf.verified_bot_category gives the bot's type, such as Search Engine Crawler or AI Search. A custom rule with the Skip action can then exempt those requests from rate limiting, Managed Rules and Super Bot Fight Mode:
In the Cloudflare dashboard, open Security rules and create a custom rule.
Set the expression to match verified bots, for example cf.client.bot, optionally narrowed by cf.verified_bot_category.
Under Choose action, select Skip and tick the products that were blocking the crawler.
Save, then re-run URL Inspection's live test on an affected page.
On another WAF, check whether it offers a managed list of verified crawlers before writing user agent matches by hand.
What does a 401 mean, and should you fix it?
A 401 means the page asked for authorization: a login wall, HTTP basic auth, or a password-protected preview. Search Console's help says that if you want Googlebot to index the page, "either remove authorization requirements for this page, or else allow Googlebot to access your pages by verifying its identity." For a page that is private on purpose, fix your links instead.
The Crawl Stats help warns about the second option: "Googlebot can be spoofed, so allowing entry for Googlebot effectively removes the security of the page." So for a genuinely private page:
Take it out of your XML sitemap.
Remove internal links to it from public pages, or accept that Google will keep finding it.
If a whole staging or preview deployment returns 401, make sure no public page or sitemap points at its hostname.
If a public page returns 401 by mistake, it is usually a password-protected deployment or an auth middleware whose matcher covers too many routes. Loading the URL in a private browser window reproduces what Googlebot sees.
Do the same rules block AI crawlers?
Usually, and with less chance of anyone noticing, because no Search Console report tells you. A firewall that challenges Googlebot tends to challenge OAI-SearchBot, PerplexityBot, Claude-SearchBot and the user-triggered fetchers too. Those are the crawlers behind AI search answers, and each company documents its user agents and publishes IP ranges.
Crawler
What it does, per its operator
IP ranges
OAI-SearchBot (OpenAI)
Surfaces websites "in search results in ChatGPT's search features"
https://openai.com/searchbot.json
ChatGPT-User (OpenAI)
Visits pages for "certain user actions in ChatGPT"
https://openai.com/chatgpt-user.json
PerplexityBot
Surfaces and links websites in Perplexity search results
https://www.perplexity.com/perplexitybot.json
Perplexity-User
Visits pages when users ask Perplexity a question
https://www.perplexity.com/perplexity-user.json
Claude-SearchBot (Anthropic)
Navigates the web "to improve search result quality for users"
Not listed on the help page read
Claude-User (Anthropic)
Accesses websites when users ask Claude questions
Not listed on the help page read
OpenAI recommends "allowing OAI-SearchBot in your site's robots.txt file and allowing requests from our published IP ranges," and Perplexity's docs say a site behind a WAF "may need to explicitly whitelist Perplexity's bots." You can test user-agent rules with curl, for example curl -s -o /dev/null -w "%{http_code}\n" -A "OAI-SearchBot/1.4" https://example.com/, but a well-configured WAF verifies IPs, so a spoofed request from your laptop may be refused even when the real crawler is allowed. The Cloudflare AI crawler guide covers Cloudflare's AI bot settings, and checking whether AI bots crawl your site shows how to confirm real visits in your logs.
If the blocked requests turn out to be 5xx rather than 403, the cause is often the same rules under load; the server error (5xx) guide covers that case.
Check what bots get from your site
To see whether bot protection answers crawlers differently from browsers, run the free scan on LaunchScaler. It takes the URL, no account needed, and runs 156 checks across 6 of its 7 categories at no cost. Its AI visibility checks request your pages as AI retrieval crawlers and flag a 403, a challenge page or an empty page where a browser gets real content, and do the same for the agent user agents ChatGPT and Claude send when a user asks them to act on a site. Its search checks cover the robots.txt and noindex side. Fix what it flags, then confirm with URL Inspection's live test, which is the only test that uses Googlebot's real IPs.
Find which layer returned the 403, usually a CDN bot rule, a WAF, a rate limit or geo-blocking, by searching its logs for the failing requests. Then allow verified crawlers by IP range or by your CDN's verified-bot field, and confirm with URL Inspection's live test.
03
What is the difference between a 401 and a 403 in Search Console?
A 401 means the page asked for authorization, such as a login wall or HTTP authentication. A 403 means the server refused access. If a 401 page is meant to be private, remove it from your sitemap and internal links rather than letting Googlebot in.
04
Can I allow Googlebot by user agent?
No. Anyone can send Googlebot's user agent string. Verify requests with a reverse and forward DNS lookup, or match the IP against the ranges Google publishes in common-crawlers.json.
05
Does a 403 also block AI crawlers?
Often, yes. The same bot protection that blocks Googlebot tends to block OAI-SearchBot, PerplexityBot and Claude's crawlers, which feed AI search answers. Each of these companies publishes IP ranges you can allow.
Google's crawl budget guide is for sites with a million pages, or 10,000 changing daily. What a small site should fix instead, and how to read Crawl Stats.
Crawled - currently not indexed means Google fetched the page and chose not to index it. Triage by URL pattern, then merge, noindex or improve each page.