Indexed, though blocked by robots.txt: why it happens and the fix
Google indexed the URL from links without reading it, because robots.txt blocks the crawl. Unblock it and add noindex to drop it, or unblock it to rank it.
LLaunchScaler·Published ·7 min read
Indexed, though blocked by robots.txt means Google added the URL to its index from links on other pages without ever reading it, because your robots.txt forbids Googlebot from fetching the page. It happens because robots.txt controls crawling, not indexing, and the fix depends on what you want: to remove the URL, lift the Disallow and add noindex; to rank it, lift the Disallow and request indexing.
The step people skip is the first one. A noindex on a page that robots.txt blocks is never seen, so adding one without unblocking the page changes nothing.
What does "Indexed, though blocked by robots.txt" mean?
It means the URL is in Google's index, but Google has not read the page. Search Console's help says: "Google always respects robots.txt, but this doesn't necessarily prevent indexing if someone else links to your page. Google won't request and crawl the page, but we can still index it, using the information from the page that links to your blocked page."
That information is the link itself: the URL and the anchor text around it. Because Google never fetched the page, "any snippet shown in Google Search results for the page will probably be very limited." Google's robots.txt introduction describes the same result: the URL "can still appear in search results, but the search result won't have a description."
Search Console lists this under "Improve page experience", the table for issues that did not stop indexing, rather than among the reasons a page was not indexed. The page indexing report guide explains that split.
Why does Google index a page it was told not to crawl?
Because the two instructions cover different steps. Robots.txt tells crawlers which URLs they may request; Google's introduction says it is used "mainly to avoid overloading your site with requests" and "is not a mechanism for keeping a web page out of Google." Indexing is a separate decision, and a URL that many pages link to can earn a place in the index on those links alone.
Common sources of those links:
Questions, answered
What people ask about this
01
What does indexed, though blocked by robots.txt mean?
Google indexed the URL without crawling it, using links from other pages, because your robots.txt stops Googlebot from fetching the page. The result usually shows little or no snippet, since Google never read the content.
02
How do I fix indexed, though blocked by robots.txt?
Your own navigation, footer or sitemap still linking to a section you blocked.
Old URLs that other sites linked to before you added the Disallow.
Parameter or filter URLs your templates link to, which you blocked to save crawl.
Staging, admin or login paths linked from a public page by mistake.
Google also warns against using robots.txt for canonicalization for this reason: it "may still index URLs that are disallowed in robots.txt without their content."
How do you find what is linking to the blocked URL?
Start with Search Console, which records where Google found each URL. In URL Inspection, the indexed result's Discovery section shows the sitemaps that list the URL and a referring page Google may have used to find it. Those two fields usually point straight at the template or file that keeps feeding Google the address.
Paste an example URL from the warning into URL Inspection.
Expand the Page indexing section and read Sitemaps and Referring page under Discovery.
If a sitemap is listed, remove the blocked URL from it. Google's help names a sitemap "that includes URLs that are blocked for crawling by robots.txt" as a cause of error spikes, and a sitemap should list only URLs you want indexed.
If the referring page is on your site, fix the link in that page's template, not just on the one page.
In the Page indexing report, set the sitemap filter to "All submitted pages" to see how many blocked URLs your sitemaps still submit.
Links from other sites are outside your control, which is why the permanent fix below works on your side of the connection: a readable noindex or a 404 overrides whatever links point at the URL.
Why doesn't adding noindex work while the page is blocked?
Because Google has to fetch the page to read the noindex, and the robots.txt rule stops the fetch. Google's noindex documentation says it directly: "If the page is blocked by a robots.txt file or the crawler can't access the page, the crawler will never see the noindex rule, and the page can still appear in search results." The two directives cancel each other out.
URL Inspection shows the trap. Its help says that when a page is blocked by robots.txt, "Indexing allowed? will always be 'Yes' because Google can't see and respect any noindex directives." So a page can carry a perfectly valid noindex and still report indexing as allowed, because Google has no way to read it.
Let Google crawl the page so it can read a noindex, then wait for the recrawl. Google's Removals help lists the permanent options: return 404 or 410, put the content behind a password, or use noindex, and it adds "Do not use robots.txt as a blocking mechanism." Order matters, because the noindex must be readable before it can work.
Add the noindex first: <meta name="robots" content="noindex"> in the page's HTML, or X-Robots-Tag: noindex in the response header for PDFs and other files. The excluded by noindex guide shows how to confirm it is served.
Remove or narrow the Disallow rule that covers the URL, so Googlebot may fetch it. Google caches robots.txt for up to 24 hours, so the change can take a day to apply; the robots.txt report in Search Console has a Request a recrawl option.
Run URL Inspection's live test. Crawl allowed? should say Yes and Indexing allowed? should now say No, naming the noindex.
Click Request indexing so Google recrawls the page and reads the noindex. Google says crawling "can take anywhere from a few days to a few weeks."
Watch the URL move from this warning into Excluded by 'noindex' tag in the Page indexing report. That move is the fix taking effect.
Only then decide whether to re-block. Keeping the noindex and the crawl open is the safer end state; if you re-block, the noindex goes unseen again, and Google says that if it "can find other information about this page without loading it," there is "a very small chance that the page might still be indexed."
If the content should not exist at all, deleting it and returning 404 or 410 is the cleaner route. Google says its indexing pipeline removes a previously indexed URL that returns 404 or 410.
What if the page should be in search?
Then the robots.txt rule is the bug. Remove the Disallow line that blocks it, or add a longer Allow rule that reopens just that path, and ask Google to crawl it. Once Google can read the page, it can index the content and show a normal snippet instead of a bare URL.
Find the rule that blocks the URL. Google applies the most specific user-agent group and, inside it, the longest matching path; the blocked by robots.txt guide walks through both.
Edit that one line, deploy, and fetch https://yoursite.com/robots.txt to confirm the new version is served.
In URL Inspection, run Test live URL and check that Crawl allowed? says Yes.
Click Request indexing.
Make sure the page has a self-referencing canonical and is linked from indexed pages, so Google treats it as a page worth indexing in full.
Unblocking also lets Google render the page. If the Disallow covered CSS or JavaScript folders the page needs, lifting it can change what Google sees across many URLs, which is usually an improvement.
Should you use the Removals tool?
Only as an emergency brake, and only alongside a permanent fix. Google's help says the Removals tool "provides only a temporary removal of about six months," after which the URL can reappear. It is the right first step when a private URL must leave results today, but it does not replace the noindex or the 404.
The tool also does not stop crawling. Google says blocking a URL there "does not prevent Google from crawling your page, only from showing it in Search results." Use it while you put the permanent fix in place:
Open Removals in Search Console, select the Temporary Removals tab, and click New request.
Choose Temporarily remove URL and enter the URL.
Apply the permanent fix above: noindex with the crawl allowed, a 404 or 410, or a password.
Confirm the fix with URL Inspection before the six months run out.
Check for a noindex hidden behind a robots rule
To see whether a page combines both directives, run the free scan on LaunchScaler. It needs only the URL, no account, and runs 156 checks across 6 of its 7 categories at no cost. One of its search checks exists for exactly this case: it flags a URL that is Disallowed in robots.txt while carrying a noindex in its meta tag or header, since the two defeat each other. It also flags a Disallow rule that blocks a page you want ranked, and a noindex in either location on a page that should be indexed. Fix what it reports, then confirm Crawl allowed? and Indexing allowed? in URL Inspection's live test.
Decide whether the URL should be in search. To remove it, lift the Disallow rule, add a noindex tag or header, and wait for Google to recrawl it. To keep it, remove the Disallow rule and request indexing.
03
Why is Google ignoring my noindex tag on a page blocked by robots.txt?
Google never sees it. Robots.txt stops Googlebot fetching the page, so the noindex in its HTML or headers is never read. Google's documentation says the page can still appear in search results in that case.
04
Does the Removals tool fix indexed, though blocked by robots.txt?
Only for about six months. Google calls the Removals tool temporary and says to make the removal permanent with a 404 or 410, a password, or a noindex tag, and not to use robots.txt as the blocking mechanism.
05
Is indexed, though blocked by robots.txt an error?
Search Console shows it as a warning in the Improve page experience table, not as a reason a page was left out. It only needs action when the URL is one you want out of search, or one you want to rank with its full content.
Most 404s in Search Console are harmless. Fix the ones your sitemap or own pages link to, and 301 URLs with backlinks to their closest equivalent page.