Blocked by robots.txt: fix the rule without unblocking the wrong pages
Blocked by robots.txt means a rule stops Googlebot fetching the URL. Find the group Googlebot obeys and the longest matching rule, then edit that line.
LLaunchScaler·Published ·7 min read
Blocked by robots.txt means a rule in your robots.txt file stops Googlebot from fetching the URL, so Google cannot read the page. To fix it without opening the whole site, find the user-agent group Googlebot obeys, find the longest rule in that group that matches the URL's path, and change that one line.
Google resolves robots.txt with two precise rules, one for groups and one for paths. Once you apply them the way Google does, the line to edit is usually obvious.
What does "Blocked by robots.txt" mean?
It means Googlebot requested your robots.txt, found a rule covering this URL, and did not fetch the page. Search Console's help, which calls the reason "URL blocked by robots.txt," says: "This page was blocked by your site's robots.txt file." It also warns that the block "does not guarantee that the page won't be indexed through some other means."
So the status only needs fixing when the listed URL is a page you want in search. Admin areas, carts, internal search results and API routes are commonly blocked on purpose. Open the row, read the example URLs, and decide which ones should be crawled. The page indexing report guide shows where this row sits among the others.
Which user-agent group does Googlebot follow?
Only the most specific group that matches it. Google's specification says "Only one group is valid for a particular crawler," found by "the group with the most specific user agent that matches the crawler's user agent. Other groups are ignored." A User-agent: Googlebot group therefore replaces the User-agent: * group for Googlebot entirely.
This catches people out constantly. Take this file:
Googlebot follows only the second group. It may crawl /admin/ and , because those rules live in a group it ignores, and Google says "user agent specific groups and global groups (*) are not combined." The reverse happens too: a group with , left from a staging setup, blocks Googlebot from everything while every other crawler reads the permissive group.
Questions, answered
What people ask about this
01
What does blocked by robots.txt mean in Search Console?
A rule in your robots.txt file stops Googlebot from fetching the URL, so Google cannot read the page. Google can still index the URL from links alone, without a snippet.
If several groups name the same crawler, Google merges their rules into one group.
The user-agent value is case-insensitive, and trailing text is ignored, so googlebot/1.2 and googlebot* both mean googlebot.
A crawler with no group of its own, such as Googlebot Storebot in Google's example, falls back to *.
Which rule wins inside the group?
The longest matching path. Google says its crawlers apply "the most specific rule based on the length of the rule path." If an Allow and a Disallow match with paths of equal length, Google "uses the least restrictive rule," so Allow wins the tie. Paths are case-sensitive and match from the start of the URL path.
URL
Rules in the group
Rule that applies
Why
/page
allow: /p, disallow: /
allow: /p
Longer path
/folder/page
allow: /folder, disallow: /folder
allow: /folder
Equal length, least restrictive wins
/page.htm
allow: /page, disallow: /*.htm
disallow: /*.htm
Longer path
/
allow: /$, disallow: /
allow: /$
Longer path
/page.htm
allow: /$, disallow: /
disallow: /
/$ only matches the root
Those rows are Google's own examples. Two wildcards are supported: * matches zero or more characters and $ marks the end of the URL. Disallow: /blog also blocks /blog-post and /blogging, because matching is by prefix, which is a common way a short rule blocks more than intended.
A worked case: a group holds Disallow: /docs and Allow: /docs/public/. For /docs/public/setup, both rules match, and the Allow path has 13 characters against 5, so the page is crawlable. For /docs/internal/keys, only the Disallow matches, so it stays blocked. For /docs-archive, the Disallow matches too, because /docs is a prefix of it, which may not be what the author wanted.
How do you fix the rule without unblocking the wrong pages?
Change the narrowest line that produces the block, and leave everything else alone. Deleting a whole Disallow line, or the group, often reopens areas that were blocked for a reason. A longer Allow rule that carves out only the path you need, or a Disallow narrowed to what you meant, fixes the listed URLs and nothing more.
Find the file that governs the URL. Robots.txt lives at the root of each protocol and host, and Search Console's help says "a given page can be affected by only one robots.txt file." For https://blog.example.com/post, that is https://blog.example.com/robots.txt, not the one on example.com.
Find the group Googlebot follows: a group naming Googlebot if one exists, otherwise *.
List every rule in that group whose path matches the URL, including wildcard rules.
Pick the winner by path length, with Allow winning ties.
Edit that line. To reopen one folder under a blocked parent, add a longer Allow: Allow: /app/docs/ beside Disallow: /app/. To stop a prefix from catching neighbours, end it with a slash or $: Disallow: /blog$ blocks only /blog.
Deploy the file and confirm it is served as plain text with a 200 at https://yoursite.com/robots.txt.
If your site runs on a hosting service where the file is hard to edit, Search Console's help points you to the host's own documentation on blocking or unblocking pages.
Why doesn't robots.txt keep a page out of Google?
Because it controls crawling, not indexing. Google's introduction says robots.txt "is not a mechanism for keeping a web page out of Google," and that a blocked URL "can still appear in search results, but the search result won't have a description." If other pages link to it, Google can index the address without ever reading it.
To keep a page out of search, do the opposite of blocking it: allow the crawl and add <meta name="robots" content="noindex"> or an X-Robots-Tag: noindex header. Google's noindex documentation is explicit that "if the page is blocked by a robots.txt file or the crawler can't access the page, the crawler will never see the noindex rule." The indexed though blocked by robots.txt guide walks through that case, and the X-Robots-Tag vs robots.txt guide compares the two tools.
Should you block CSS and JavaScript files?
Not the ones a page needs to render. Google renders pages before indexing them, and its robots.txt introduction says that if the absence of resources makes "the page harder for Google's crawler to understand the page, don't block them." Blocking /static/, /_next/, /assets/ or /wp-includes/ can leave Google looking at a broken or empty page.
Google's own sample file shows the pattern for a shared folder: block it for everyone, then allow Googlebot back in.
The Crawl Stats help makes the same point from the other side: when crawling drops after a new robots.txt rule, "if Google needs specific resources such as CSS or JavaScript to understand the content, be sure you do not block them from Googlebot." To check a page, run URL Inspection's live test, open View tested page, and look at the screenshot and the list of page resources that could not be loaded.
How do you check the fix and get Google to recrawl?
Test the URL with Google's live test, then refresh Google's copy of the file. Google "generally caches the contents of robots.txt file for up to 24 hours," so without a nudge the old rules can apply for up to a day after you deploy. Search Console's robots.txt report lets you ask for a faster refresh.
In URL Inspection, paste the URL and click Test live URL. Crawl allowed? should say Yes.
Open Search Console's robots.txt report. It shows the files Google found for your top 20 hosts, when each was last fetched, and any parsing issues. It works only for Domain properties or URL-prefix properties without a path.
Click the file to see the version Google last fetched, and Versions to see the fetches from the last 30 days.
Open the menu next to the file and click Request a recrawl. The help says this suits a change "to unblock some important URLs," though it "doesn't guarantee an immediate recrawl of unblocked URLs."
Request indexing for your most important unblocked pages, then click Validate fix on the Blocked by robots.txt row for the rest.
Watch the robots.txt fetch status as well. Google treats a 4xx for robots.txt (except 429) as "no crawl restrictions," but a 5xx stops crawling of the whole site for the first 12 hours, then Google falls back to the last good version for up to 30 days. A robots.txt that times out can block more than any rule in it. The robots.txt example for SaaS gives a starting file that avoids the common mistakes above.
Check which rule applies before you deploy
To confirm how a robots rule treats a page before Google recrawls, run the free scan on LaunchScaler. Give it the URL, no account needed, and it runs 156 checks across 6 of its 7 categories at no cost. Its search checks flag a robots.txt Disallow that blocks the page you gave it under the Googlebot or * group, a noindex that sits behind a robots.txt block where Google cannot see it, and a missing Sitemap: line. Its AI visibility checks read the same file for the answer-engine crawlers, flagging rules that block the retrieval bots behind ChatGPT, Perplexity and Claude citations, or a wildcard that catches them along with training bots.
Find the user-agent group Googlebot follows, find the longest rule in it that matches the URL's path, and narrow that Disallow or add a longer Allow. Then check Crawl allowed in URL Inspection's live test and request a recrawl of robots.txt.
03
Does robots.txt stop a page from being indexed?
No. Google says robots.txt is not a mechanism for keeping a web page out of Google. To keep a page out of search, allow crawling and add a noindex meta tag or X-Robots-Tag header.
04
How long does Google take to see robots.txt changes?
Google generally caches robots.txt for up to 24 hours. You can ask for a faster refresh with Request a recrawl in Search Console's robots.txt report.
05
If I have a Googlebot group, does Googlebot still follow the User-agent: * rules?
No. Google's crawlers obey only the most specific group that matches them and ignore the others, and Google does not combine a user-agent-specific group with the * group.
A 403 to Googlebot usually comes from bot protection, not your app. Find the layer in your WAF logs and allow verified crawlers by IP, never by user agent.
Google's crawl budget guide is for sites with a million pages, or 10,000 changing daily. What a small site should fix instead, and how to read Crawl Stats.