X-Robots-Tag vs robots.txt vs meta noindex: which one to use
robots.txt controls crawling; meta robots and X-Robots-Tag control indexing. Which to use for pages, PDFs and folders, plus Nginx, Vercel and Next.js.
LLaunchScaler·Published ·8 min read
robots.txt controls crawling: whether a crawler may fetch a URL at all. The robots meta tag and the X-Robots-Tag header control indexing: whether a URL Google has fetched may appear in results. Use robots.txt to keep crawlers out of private or endless paths, the meta tag to noindex an HTML page, and the X-Robots-Tag header for PDFs, images and whole folders.
The three are not interchangeable, and combining them the wrong way can leave a page you meant to hide sitting in Google's results.
What is the difference between X-Robots-Tag, robots.txt and meta robots?
robots.txt is one text file per host that lists paths crawlers should not fetch. The robots meta tag is a line in a page's HTML that says how to index and show that page. X-Robots-Tag is an HTTP response header that carries the same rules as the meta tag, for any file type. Only the last two can keep a URL out of results.
robots.txt
Meta robots tag
X-Robots-Tag header
Where it lives
/robots.txt at the root of each host
<meta name="robots"> in the page's HTML
The HTTP response for the URL
What it controls
Crawling (fetching)
Indexing and snippets
Indexing and snippets
Questions, answered
What people ask about this
01
What is the difference between X-Robots-Tag and robots.txt?
robots.txt tells crawlers which URLs they may fetch. X-Robots-Tag is an HTTP response header that tells search engines what to do with a URL once fetched, such as noindex. A URL blocked in robots.txt is never fetched, so its X-Robots-Tag is never read.
02
When should I use X-Robots-Tag instead of a meta robots tag?
One response, or every response a server rule matches
When Google reads it
Before fetching any URL
After fetching and parsing the page
After fetching, from the headers
Keeps a URL out of results?
No, a blocked URL can still be indexed from links
Yes, with noindex
Yes, with noindex
Example
Disallow: /app/
<meta name="robots" content="noindex">
X-Robots-Tag: noindex
Google's own robots.txt introduction is direct about the first column: robots.txt "is not a mechanism for keeping a web page out of Google." A disallowed URL can still be indexed if other pages link to it, and it then shows in results without a description.
The meta tag and the header are equivalent for HTML. Google's specification says any rule that can be used in a robots meta tag can also be specified as an X-Robots-Tag, and that you should pick whichever is more convenient for your stack and the content type.
Which one should you use for each job?
Pick by what you want to happen and what kind of file it is. To keep a page out of results, use noindex, and leave the page crawlable. To stop crawlers wasting requests on URLs nobody should land on, use robots.txt. For anything that is not HTML, the header is the only indexing control you have.
Goal
Use
Why
Keep an HTML page out of Google
Meta noindex (or the header)
Works page by page; Google drops the page once it recrawls and sees it.
Keep a PDF, image or video out of Google
X-Robots-Tag: noindex
A PDF has no <head> to put a meta tag in.
Noindex a whole folder or file type
X-Robots-Tag set in server or host config
One rule covers every matching response.
Stop crawling of /api/, /app/ or filter URLs
robots.txt Disallow
Saves requests; these URLs have no value to a searcher.
Remove a page that is already indexed
Meta or header noindex, with crawling allowed
Google must fetch the page again to see the rule.
Hide a staging site
HTTP authentication
robots.txt cannot enforce anything, and Google recommends password protection for private content.
For a normal SaaS or content site, that works out to a short robots.txt for the app and API, noindex in the metadata of the few HTML pages you want hidden, and an X-Robots-Tag rule for any downloads folder.
Why must a noindex page be crawlable?
Google only sees a noindex when it fetches the page. Its documentation says that for noindex to be effective, "the page or resource must not be blocked by a robots.txt file." If robots.txt blocks the URL, Google never reads the meta tag or the header, and the page can keep appearing in results if other pages link to it.
This is the combination that produces the Search Console status "Indexed, though blocked by robots.txt", covered in detail in the indexed though blocked by robots.txt guide. URL Inspection makes the trap visible: when "Crawl allowed?" is No, "Indexing allowed?" always shows Yes, because Google cannot see any noindex behind the block.
The fix is counterintuitive but simple. To get a blocked page out of Google:
Remove the Disallow rule that covers it.
Make sure the page returns noindex in the meta tag or the header.
Wait for Google to recrawl it, or request indexing on the URL so it is fetched sooner.
Only after it has dropped out, and only if you also want to save crawl requests, add the Disallow back.
The reverse mistake happens too: a page you want indexed shows up under "Blocked by robots.txt" in the Page indexing report. The blocked by robots.txt guide covers finding and removing the rule.
How do you set X-Robots-Tag on Nginx, Vercel and Next.js?
Set the header in whatever layer answers the request: the web server, the host's config file or the framework. Each takes the header name X-Robots-Tag and a value such as noindex or noindex, nofollow, scoped to a path or file pattern. Then confirm it with curl -I, because the header never appears in the page source.
Nginx, for every PDF on the site (Google's own example uses the same location block):
Two Nginx details decide whether it works. Without always, add_header only adds the header to responses with status 200, 201, 204, 206, 301, 302, 303, 304, 307 or 308. And add_header directives are inherited from the parent level only if the current level defines none, so an add_header for another header inside the same location silently drops every header set higher up.
Apache, in .htaccess or httpd.conf:
<Files ~ "\.pdf$">
Header set X-Robots-Tag "noindex, nofollow"
</Files>
For an HTML page in the Next.js App Router, the meta tag is usually simpler: export const metadata = { robots: { index: false, follow: true } } in the page renders <meta name="robots" content="noindex, follow">.
The header can also name one crawler: X-Robots-Tag: googlebot: noindex applies only to Google, and a rule with no user agent applies to all crawlers.
What is the difference between noindex and nofollow?
noindex asks search engines not to show a page in results. nofollow asks them not to follow links. As a page-level rule in the meta tag or header, nofollow covers every link on the page. As rel="nofollow" on a single <a> tag, it covers only that link. They are independent, and none is shorthand for both.
Directive
Where
What it asks
noindex
Meta tag or header
Don't show this page, file or resource in results.
nofollow
Meta tag or header
Don't follow any link on this page.
none
Meta tag or header
Same as noindex, nofollow.
rel="nofollow"
One <a> element
Don't associate the site with this link or crawl the target from it.
rel="sponsored", rel="ugc"
One <a> element
Paid links, and user-generated links such as comments.
Google says links marked with these values will generally not be followed, and that the target can still be crawled if it is found another way, such as a sitemap or a link from another site. For links inside your own site, Google's guidance is to use a robots.txt Disallow instead of rel="nofollow" when you don't want a URL fetched.
The page-level combination noindex, follow asks Google to drop the page but still crawl its links. How long Google keeps following links on a page it has stopped indexing is a separate question, answered in the noindex, follow guide.
What happens when the directives conflict?
When robots rules conflict, Google applies the more restrictive one. A page that says index in its meta tag and noindex in its X-Robots-Tag header is treated as noindex, and max-snippet:50 alongside nosnippet means no snippet.
Three more conflicts come from how and when rules are read:
A robots.txt block beats everything else, because Google never fetches the page to read the other rules.
A noindex in the original HTML can stop Google rendering the page. Google's JavaScript guide says that when it encounters the noindex tag it may skip rendering, so JavaScript that later removes the noindex may not work. If you want the page indexed, don't ship noindex in the initial HTML.
A meta robots tag in the <body> still counts. Google says it respects robots meta tags outside the head, so a stray noindex injected by a widget anywhere in the page could deindex it.
When a page drops out and you cannot see why, open URL Inspection and read "Indexing allowed?". It names the source, meta tag or header, and the excluded by noindex tag guide walks through tracing it to the template or config that adds it.
Find the directives your pages actually send
Header noindex rules are invisible in a browser and easy to ship by accident through a host setting or a middleware. Run the free scan on LaunchScaler to check the live responses without an account: it runs 156 checks across 6 of its 7 categories at no cost, and its search checks read noindex in both the meta robots tag and the X-Robots-Tag header, flag a noindex sitting behind a robots.txt block that Google can never see, and flag robots.txt rules that block a page you want ranked.
Use X-Robots-Tag for anything that is not HTML, such as PDFs, images and video files, and when you want one server rule to cover a whole folder or file type. For a normal HTML page, the meta tag and the header have the same effect.
03
Does noindex work if the page is blocked in robots.txt?
No. Google has to crawl the page to see the noindex, so a robots.txt block hides it. The URL can then stay indexed from links, without a description.
04
What is the difference between noindex and nofollow?
noindex asks search engines not to show the page in results. nofollow in a robots meta tag or header asks them not to follow any link on the page, while rel="nofollow" on one link applies only to that link.
05
What happens if the meta tag and the X-Robots-Tag header disagree?
Google applies the more restrictive rule. A page with index in the meta tag and noindex in the header is treated as noindex.
Alternate page with proper canonical tag means Google honoured your canonical and indexed that URL. It only hurts when a template canonicalises every page.
Blocked by robots.txt means a rule stops Googlebot fetching the URL. Find the group Googlebot obeys and the longest matching rule, then edit that line.
A 403 to Googlebot usually comes from bot protection, not your app. Find the layer in your WAF logs and allow verified crawlers by IP, never by user agent.