robots.txt and noindex checkers: tools that find what blocks a page
Six robots.txt and noindex checkers compared: LaunchScaler, Search Console's report and URL Inspection, Bing's tester, an extension and Google's parser.
LLaunchScaler·Published ·7 min read
The robots.txt and noindex checkers worth using are LaunchScaler's free scan, Search Console's robots.txt report, Search Console's URL Inspection, Bing Webmaster Tools' robots.txt tester, the Robots Exclusion Checker extension for Chrome, and Google's open-source robots.txt parser. Each one misses something, so the useful question is which gap matters for your page.
The gaps fall into three groups: some tools only work on sites you have verified, some read the noindex meta tag but not the X-Robots-Tag header, and most test robots.txt rules without checking whether a firewall lets the crawler in at all. Each tool's terms were checked on its own pages in September 2026.
What can block a page, and which checker sees it?
A page can be blocked from crawling by robots.txt or a firewall, and from indexing by noindex in the HTML meta tag or in the X-Robots-Tag response header. robots.txt is one file per host; noindex lives on each page. A checker that reads only one of these places can report a page as fine when it isn't.
Checker
Works on
robots.txt per crawler
Meta noindex
X-Robots-Tag noindex
AI bots
Firewall blocks
1. LaunchScaler free scan
Any public URL, no account
Yes
Yes
Yes
OAI-SearchBot, Claude-SearchBot, PerplexityBot
Questions, answered
What people ask about this
01
How do I check if robots.txt is blocking a page?
For a site you own, inspect the URL in Google Search Console and read 'Crawl allowed?'. For any site, run LaunchScaler's free scan or use a robots.txt tester that matches the URL against the file for a chosen crawler, such as Googlebot or OAI-SearchBot.
The order here is LaunchScaler, the robots.txt report, URL Inspection, Bing's tester, Robots Exclusion Checker, then Google's parser. LaunchScaler leads because it is the only one that checks both robots.txt and both forms of noindex on any URL without an account or an install, and also tests whether the AI retrieval bots get through the firewall. The rest are ordered from most authoritative for your own site to most specialised.
1. LaunchScaler: robots rules per crawler and both noindex forms, on any URL
LaunchScaler's readiness scan takes any live URL, needs no account, and runs 156 checks across 6 of its 7 categories at no cost. Its search and AI visibility checks read robots.txt the way each crawler does and look for noindex in both the meta tag and the response header.
The checks that cover blocking:
a robots.txt Disallow that matches a page you want ranked, for the Googlebot group or *
rules that block the AI retrieval bots OAI-SearchBot, Claude-SearchBot or PerplexityBot, including a blanket Disallow: / that catches them along with training bots
a firewall or bot-protection rule that returns 403 or a challenge page to those retrieval bots while browsers get a 200
noindex in the meta robots tag or the X-Robots-Tag header on a page meant to rank
a noindex that sits behind a robots.txt block, where Google can never read it
a missing sitemap, or a robots.txt with no Sitemap: line
The retrieval-bot checks matter because the bots have separate names from the training bots. OpenAI says sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers," and Perplexity recommends allowing PerplexityBot to appear in its results. A file that blocks GPTBot to opt out of training is fine; one that blocks OAI-SearchBot too is not. The robots.txt for AI bots guide gives the per-bot groups.
What it doesn't show is Google's own record of your site. Use Search Console for that.
2. Search Console robots.txt report: the file Google actually fetched
The robots.txt report shows which robots.txt files Google found for the top 20 hosts on your site, when each was last fetched, its size, and any parse warnings or errors, with the last fetched version and 30 days of earlier versions. It is free and authoritative, because it shows Google's copy, not yours.
Open it from Settings, then robots.txt. It is available only for domain-level properties: a Domain property, or a URL-prefix property without a path. A fetch status of "Not Fetched - Not found (404)" is fine and means Google crawls everything; any other failure deserves a look. After an urgent fix, open the file's menu and choose Request a recrawl.
The report doesn't test a URL against the rules or read any noindex. Google's help points to URL Inspection for testing whether a specific URL is blocked, and to its open-source parser for developers.
3. Search Console URL Inspection: both directives for one URL on your site
URL Inspection answers both questions for one URL: "Crawl allowed?" says whether a robots.txt rule blocks Googlebot, and "Indexing allowed?" says whether Google found a noindex, naming the meta tag or the X-Robots-Tag header as the source. Test live URL checks the current page. It only works on properties you have verified, with a daily inspection limit.
It is the tool that settles disputes with extensions, because it reports the header as well as the tag. It also exposes the combination other tools miss: when robots.txt blocks a page, "Indexing allowed?" always shows Yes, because Google can't see any noindex behind the block. The blocked by robots.txt guide and the excluded by noindex tag guide cover what to do with each result.
4. Bing Webmaster Tools robots.txt tester: rules as Bingbot reads them
Bing's tester checks a URL against your robots.txt and shows which statement blocks it and for which user agent, with a toggle between Bingbot and AdIdxbot. It includes an editor: change the rules, test again instantly, then download the file to upload, and use Fetch latest after you publish it.
It is free inside Bing Webmaster Tools, for sites you have added there. Bing's index also feeds Microsoft Copilot, so a rule that blocks Bingbot matters for AI answers as well as Bing search. It doesn't read noindex.
5. Robots Exclusion Checker: an extension that reads the header too
Robots Exclusion Checker is a Chrome extension that flags, for the page you're viewing, the robots.txt rule that applies, meta robots directives, the X-Robots-Tag header, canonical mismatches, and nofollow, ugc and sponsored links. You can switch its user agent between Googlebot, Googlebot News, Bing and Yahoo, and since version 1.3.0 it monitors robots.txt rules for 14 AI bots across six companies.
This is where the extension gap shows up. Its own listing notes that many extensions don't flag robots.txt blocks, and simpler ones, such as SeeRobots, describe themselves as displaying a site's meta robots information, which leaves a header-only noindex invisible. Before you trust a clean result from any extension, check that its description mentions X-Robots-Tag, or confirm with curl -sI https://example.com/page | grep -i x-robots-tag.
Its AI bot feature reports what robots.txt says for each bot. The listing doesn't describe requesting the page as those bots, so a firewall rule that challenges a crawler the file allows won't necessarily show up.
6. Google's open-source robots.txt parser: the exact matching logic, locally
Google publishes its robots.txt parser and matcher as a C++ library on GitHub, and its documentation describes it as the one used in Google Search. The included robots_main binary checks one URL and one user agent against a robots.txt file and prints ALLOWED or DISALLOWED.
bazel run robots_main -- ~/path/to/robots.txt Googlebot https://example.com/app/settings
It is free, runs offline, and accepts any user agent name, so you can test OAI-SearchBot or PerplexityBot against the same file. It needs Bazel or CMake to build, and it only answers the robots.txt question, which makes it a developer's tool for testing a file before you deploy it.
Which checker should you use?
Pick by the question and whose site it is. For your own site, Search Console is authoritative for Google, and Bing Webmaster Tools for Bing. For any site, or to check AI bots and firewalls, use a tool that fetches the page itself. For testing a file before deploy, use Google's parser.
Your situation
Use
A page you own dropped out of Google
URL Inspection, then the robots.txt report
You changed robots.txt and want Google to reread it
robots.txt report, Request a recrawl
You want to know if ChatGPT, Claude and Perplexity can reach you
LaunchScaler free scan
An extension says indexable but Search Console says noindex
curl -I for the X-Robots-Tag header, or URL Inspection
You're editing robots.txt and want to test before deploy
Google's parser, or Bing's tester editor
You're auditing a competitor or client site you can't verify
LaunchScaler free scan, Robots Exclusion Checker
For a page-by-page view of every indexability directive, including canonicals, redirects and soft 404s, the indexability checker comparison covers the wider set of tools.
Check robots.txt and noindex on your site now
Run the free scan on the page you're worried about. It needs only the address, no account, and reports in one pass whether robots.txt blocks it for Googlebot or the AI retrieval bots, whether a firewall challenges those bots, and whether a noindex sits in the meta tag or the X-Robots-Tag header.
Check both places a noindex can live: the robots meta tag in the HTML and the X-Robots-Tag HTTP response header. View source only shows the first; run curl -I on the URL, or use a checker that reads response headers, to see the second.
03
Does Google still have a robots.txt tester?
Google's current documentation offers two options: the robots.txt report in Search Console, which shows the file Google fetched for your top hosts, when, and any parse errors, and Google's open-source robots.txt library for testing locally. To test one URL on your site, use URL Inspection.
04
Can a robots.txt checker tell me if AI bots can crawl my site?
Some can. LaunchScaler's free scan tests the rules for OAI-SearchBot, Claude-SearchBot and PerplexityBot, and the Robots Exclusion Checker extension monitors AI bots too. A robots.txt Allow is not enough on its own if a firewall challenges those bots.
05
Why does my browser extension say a page is indexable when Google says noindex?
The noindex is probably in the X-Robots-Tag response header, not the HTML. Extensions that only read the page's meta tags miss it, while Search Console's 'Indexing allowed?' line names the header as the source.
Server error (5xx) means Googlebot got a 500-level status or a timeout. Find when it happened in Crawl Stats, match it to your logs, then fix the cause.
Couldn't fetch can be transient, but a sitemap that stays unread is usually blocked, redirected, not XML or over 50,000 URLs. Check each cause with curl.