Excluded by 'noindex' tag: find where the noindex comes from
Excluded by noindex tag means Google saw noindex in a robots meta tag or an X-Robots-Tag header. Find which one with curl and view-source, then remove it.
LLaunchScaler·Published ·7 min read
Excluded by 'noindex' tag means Google fetched the page and found a noindex rule, so it left the page out of its index. That rule lives in one of two places: a <meta name="robots" content="noindex"> tag in the HTML, or an X-Robots-Tag: noindex HTTP response header, and you fix it by finding which one your page sends and removing it at its source.
The header is the one people miss, because it never shows up in the page source. One curl command shows it.
What does "Excluded by 'noindex' tag" mean?
It means Google obeyed an instruction you sent. Search Console's help, which labels this reason "URL marked 'noindex'", says: "When Google tried to index the page it encountered a 'noindex' directive and therefore did not index it. If you do not want this page indexed, congratulations! If you do want this page to be indexed, you should remove the 'noindex' directive."
So the only question is whether each listed URL should be in search. Login pages, thank-you pages, internal search results and thin archives belong here on purpose. Your homepage, pricing page, product pages and articles do not. Open the row, read the example URLs, and treat any page you want found as a bug to trace. The page indexing report guide covers how this row fits with the others.
Where can a noindex come from?
From the HTML or from the HTTP response. Google supports both: a robots meta tag in the page, and the X-Robots-Tag header, which "can be used as an element of the HTTP header response for a given URL." Any rule valid in the meta tag works in the header, so a page can be noindexed without a single character of it in the HTML.
Source
What it looks like
Questions, answered
What people ask about this
01
What does excluded by noindex tag mean in Search Console?
Google fetched the page and found a noindex rule, either in a robots meta tag in the HTML or in an X-Robots-Tag HTTP header, so it did not index the page. If you meant to hide the page, nothing needs fixing.
CMS setting, SEO plugin, theme, framework metadata
Googlebot-only meta tag
<meta name="googlebot" content="noindex">
Same, when someone targeted Google alone
HTTP header
X-Robots-Tag: noindex
Server config, CDN rule, framework header config, hosting platform
Header for one crawler
X-Robots-Tag: googlebot: noindex
Same, with a user agent before the rule
Three details from Google's specification matter when you hunt for it. Google "will respect robots meta tags in the body section" too, not only in the <head>. When rules conflict, "the more restrictive rule applies," so a page with both index and noindex is noindexed. And the header name and values are not case sensitive, so search for x-robots-tag in any casing.
If the listed URLs are PDFs, images or other files rather than HTML pages, the header is the only place the noindex can be. Google's specification says that "to block indexing of non-HTML resources, such as PDF files, video files, or image files, use the X-Robots-Tag response header instead." So a row full of .pdf URLs traces back to whatever sets response headers for those files, typically a server, CDN or hosting rule scoped by file type or folder, and that rule is what you edit.
How do you find which one your page sends?
Check the header with curl, then the raw HTML with view-source or curl again. Do not trust your browser's DevTools Elements panel for this: it shows the live DOM after JavaScript ran, and Google says that when it meets a noindex it "may skip rendering and JavaScript execution," so a script that removes the tag later does not help.
Any output means the header is set. -I sends a HEAD request, and a few servers answer HEAD differently from GET, so if it prints nothing, confirm with a real GET that discards the body:
Or open view-source:https://example.com/pricing in Chrome and search for noindex. View-source shows the HTML the server sent, which is what Google reads first.
Confirm with Google's own fetch. In URL Inspection, click Test live URL. The Indexing allowed? line says whether a noindex was found and where. Click View tested page, then More info, to read the HTTP response headers Google received.
If curl shows no noindex and URL Inspection still reports one, check whether your server or CDN answers Googlebot differently, for example a bot rule that serves a different response, and read the tested page's headers again.
How do you remove a noindex in WordPress?
Start with the site-wide switch. In the dashboard go to Settings, then Reading, and find Search engine visibility. If "Discourage search engines from indexing this site" is ticked, WordPress prints <meta name='robots' content='noindex,nofollow' /> into the head of every page (since version 5.3, when the theme calls wp_head). Untick it and save.
The setting is easy to carry over when a staging copy becomes the live site. After that:
Open the affected post in the editor and check your SEO plugin's panel for a per-post option that excludes it from search.
Check the plugin's settings for the whole content type and for archives, tags, categories, authors and media pages, which are often set to noindex in bulk.
Purge every cache: the caching plugin, the host's page cache and your CDN. A cached page keeps serving the old meta tag after you change the setting.
Run the curl commands above against the live URL to confirm the tag is gone.
Look in three places: the robots field in your metadata, a header rule in next.config.js or your proxy, and your host. In the App Router, robots: { index: false } in a layout's metadata renders a noindex meta tag, and nested metadata fields set in a layout are inherited by every page below it that does not set its own.
Search the app for robots: and googleBot: in metadata exports and generateMetadata. A robots: { index: false } added to app/layout.tsx before launch noindexes the whole site until someone removes it.
Search next.config.js for X-Robots-Tag in headers(), and your proxy or middleware for code that sets the header. Check the condition around it: Next.js sets NODE_ENV to production for every command except next dev, so a staging build reads production too and that check cannot tell staging from live.
Check the host. Vercel adds X-Robots-Tag: noindex to every Preview Deployment automatically, and to a previous production deployment once a newer one is promoted. It omits the header when a custom domain is assigned to a non-production branch. Your production domain should never carry it; if it does, the header comes from your own config.
Vercel's own recipe for a custom-domain preview adds the header only when process.env.VERCEL_ENV === 'preview', which is the kind of explicit condition that keeps it off production.
Why does a staging noindex end up on production?
Because staging and production usually share one codebase, one CMS database or one CDN configuration, and the noindex follows whichever of them was copied. When this status appears across a whole site right after a launch or a migration, a staging setting that followed the code or data to production is the first thing to check.
The usual paths:
A pre-launch robots: { index: false } or "Discourage search engines" setting that nobody switched off on launch day.
A staging database or CMS export imported into production with its visibility setting intact.
A header rule keyed on the wrong environment variable, so it fires in production too.
A CDN or server rule written for the staging hostname that matches every hostname.
A page cache that stored a noindexed page before the fix and keeps serving it.
Remove the rule at its source, confirm it is gone from the live response, then ask Google to recrawl. Google only learns the noindex is gone when it fetches the page again, so until then the page stays out, and Google says crawling "can take anywhere from a few days to a few weeks."
Deploy the fix and purge caches.
Re-run both curl checks. Neither should print a noindex.
Make sure the page is not blocked in robots.txt. If it is, Google cannot fetch it to see the change, and URL Inspection's help says that for a robots-blocked page "Indexing allowed? will always be 'Yes'" because Google cannot see noindex rules at all. The X-Robots-Tag vs robots.txt guide explains how the two interact.
In URL Inspection, click Test live URL and confirm Indexing allowed? says Yes.
Click Request indexing for your key pages, once each; the quota is daily and repeat requests do not speed it up.
For the rest, open the Excluded by 'noindex' tag row and click Validate fix.
Check every page template for a stray noindex
To catch a noindex before Google does, run the free scan on LaunchScaler. It takes a URL, no account needed, and runs 156 checks across 6 of its 7 categories at no cost. It checks both places a noindex can hide, the <meta name="robots"> tag and the X-Robots-Tag response header, and flags a noindex that sits behind a robots.txt block where Google cannot see it, plus robots directives that contradict what the page returns. It reads the live response, so run it on your homepage and one page per template after every launch or migration.
Find where the noindex comes from with curl -I for the header and view-source for the meta tag, remove it at that source, then run Test live URL in URL Inspection and click Request indexing once Indexing allowed shows Yes.
03
How do I fix excluded by noindex tag in WordPress?
Go to Settings, then Reading, and untick Discourage search engines from indexing this site, which prints a noindex, nofollow robots meta tag on every page. Then check your SEO plugin's indexing settings for the post, its post type and its archives.
04
Why can't I see the noindex tag in my page?
It is probably in the X-Robots-Tag response header, which never appears in the HTML. Run curl -I on the URL, or open View tested page in URL Inspection and read the HTTP response headers.
05
How long after removing noindex will Google index the page?
Google has to recrawl the page first. Requesting indexing in URL Inspection queues it, and Google says crawling can take anywhere from a few days to a few weeks.
Google treats rel=canonical as a hint and weighs redirects, sitemaps, links, HTTPS and hreflang. Make every signal name one URL and Google usually follows.