Noindex on your production site: catch it before Google drops you
A noindex reaches production through a staging env var, WordPress's Discourage search engines box or a host header. Find it, remove it and recover.
LLaunchScaler·Published ·7 min read
A noindex reaches production by accident through one of three routes: a staging environment variable or build flag, the WordPress "Discourage search engines from indexing this site" box left ticked, or an X-Robots-Tag header added by the server, host or CDN. Google drops each affected page the next time Googlebot crawls it, so check both the meta tag and the response header after every deploy, and if it has already shipped, remove it, request indexing for your key pages and resubmit your sitemap.
The damage is proportional to how long the noindex stays live, because every crawl of an affected page during that window removes it from results.
How does a noindex end up on a production site?
A noindex ends up in production when the mechanism that keeps staging out of search is a setting rather than access control, and production inherits the staging value. The setting lives in one of three places (application code, the CMS, or the server and CDN layer), and each one needs a different check.
Source
Where it lives
What reaches the browser
How to spot it
Staging environment variable or build flag
Framework metadata, a layout, or a robots route
<meta name="robots" content="noindex"> on every page
View source on any page
WordPress "Discourage search engines from indexing this site"
Settings, Reading, Search engine visibility
Questions, answered
What people ask about this
01
How do I know if my production site has a noindex?
Run curl -sI on a page and look for an X-Robots-Tag header containing noindex, then search the page source for a meta tag named robots or googlebot with noindex or none in its content. In Search Console, URL Inspection's live test reports it under Indexing allowed.
02
How quickly does Google remove pages with a noindex?
<meta name='robots' content='noindex,nofollow' /> in every page head
View source, or check the box in the admin
SEO plugin or CMS setting per content type
Plugin settings for posts, pages or products
A noindex meta tag on one content type only
View source on one page of each type
Server config
nginx add_header, Apache Header set, .htaccess
X-Robots-Tag: noindex header
curl -sI on the page
Host or CDN header rule
vercel.json, netlify.toml, _headers, a CDN response rule
X-Robots-Tag: noindex header
curl -sI on the page
A staging environment variable
The usual code bug is a condition that reads an environment variable production never set. A layout that emits noindex "unless DEPLOY_ENV is production" does exactly that on the first deploy to a new host, a new project or a new region where the variable was not copied. The same pattern breaks generated robots.txt routes.
Platform defaults add a second trap. Vercel adds X-Robots-Tag: noindex to every preview deployment automatically, which teams sometimes rely on instead of writing their own rule. Its documentation also says Vercel "omits X-Robots-Tag: noindex when a custom domain is assigned to a non-production branch," so the protection you assumed on a staging domain may not be there, and a copied header rule meant to replace it may end up on production.
The WordPress Discourage search engines box
In WordPress, Settings, Reading has a Search engine visibility option labelled "Discourage search engines from indexing this site." Since WordPress 5.3, ticking it outputs <meta name='robots' content='noindex,nofollow' /> in the head of every page; up to 5.2 it changed robots.txt to Disallow: / instead. The box is often ticked during a build and forgotten at launch, or ticked again when a site is migrated from a staging copy whose database had it on.
A header added at the host or CDN
A header rule is the hardest source to see, because it never appears in the page source. Google's documentation shows how simple the rule is: in nginx, add_header X-Robots-Tag "noindex, nofollow"; and in Apache, Header set X-Robots-Tag "noindex, nofollow". A rule written for a staging hostname that matches every host, or a location block that matches / instead of /staging, puts the header on production.
How fast does Google drop a page after a noindex?
Google drops the page the next time Googlebot crawls it. Its documentation says: "Google will drop that page entirely from Google Search results, regardless of whether other sites link to it." There is no grace period and no warning first. How soon that happens depends on how often Google crawls each page, which Google does not publish.
In practice, the pages Google crawls most often are dropped first. Google also notes the reverse case: "If a page is still appearing in results, it's probably because we haven't crawled the page since you added the noindex rule." A page still in results is not proof the noindex is harmless; it may just not have been recrawled yet.
In Search Console, affected URLs move to the "Excluded by 'noindex' tag" row of the Page indexing report. That row is where you confirm the scale of the damage after the fact, and the excluded by noindex tag guide explains how to read it.
How do you check for a noindex after every deploy?
Check both places a noindex can live, the HTML and the HTTP response headers, on one URL for each page template, right after each deploy finishes. A tag check alone misses header rules, and a header check alone misses the CMS and framework settings. Both checks take seconds with curl.
Check the header: curl -sI https://yoursite.com/ | grep -i x-robots-tag. On a page you want indexed, this should print nothing or a value without noindex or none.
Check the HTML: curl -s https://yoursite.com/ | grep -io '<meta[^>]*name="\(robots\|googlebot\)"[^>]*>'. Read the content value of anything it prints.
Repeat both for one URL of each template: home, a product or pricing page, a blog post, a docs page.
Treat none as noindex. Google defines it as "equivalent to noindex, nofollow."
Check a googlebot meta tag as well as robots. Google applies both, and "in the case of conflicting robots rules, the more restrictive rule applies."
For a page you care about, run URL Inspection's Test live URL in Search Console and read "Indexing allowed?".
A noindex that JavaScript removes after load does not help. Google's JavaScript guidance says "when Google encounters the noindex tag, it may skip rendering and JavaScript execution," so the server HTML is the version that decides.
Keep robots.txt out of the fix. If a page is blocked in robots.txt, Google's documentation says "the crawler will never see the noindex rule," so blocking a noindexed page leaves the old state in place. This is also the mistake to avoid with staging copies, covered in staging site indexed by Google.
How do you recover after a noindex shipped to production?
Remove the noindex at its source, prove the fix is live, then ask Google to recrawl the pages that matter most. Google has to recrawl each page to see the change, and it says crawling "can take anywhere from a few days to a few weeks," so start with the pages that bring the most traffic.
Find every source. Check the header and the HTML on every template, because a site that had one noindex source often has two (a CMS setting and a header rule, for example).
Remove it and deploy. Clear any CDN cache that may still serve the old HTML or headers.
Verify with curl and with URL Inspection's Test live URL on several pages. "Indexing allowed?" should no longer report a noindex.
Click Request indexing for your most important URLs. Google notes "there's a quota for submitting individual URLs and requesting a recrawl multiple times for the same URL won't get it crawled any faster."
Resubmit your XML sitemap in the Sitemaps report so Google sees the full list of URLs to recrawl.
Open the Excluded by 'noindex' tag row in the Page indexing report and click Validate fix, then follow the count down over the following weeks.
A recrawled page without the noindex is eligible to be indexed again, but Google does not promise when; its recrawl documentation says requesting a crawl "does not guarantee that inclusion in search results will happen instantly or even at all." For more on the order of checks after a drop, see SEO regressions after a deploy.
How do you stop a noindex shipping again?
Make the safe state the default and test it where it cannot be skipped. Staging should be kept private by access control, and production should be indexable unless a rule explicitly and narrowly says otherwise. A check in the deploy pipeline turns the next leak into a failed build instead of a lost month.
Protect staging and preview environments with a password or an IP allowlist instead of relying on noindex. A page nobody can load cannot be indexed, and there is no setting to copy into production.
Invert risky conditions. Emit noindex only when the environment is positively identified as staging, never whenever it is "not production", so a missing variable fails open to indexable.
Add a post-deploy test that requests key URLs and fails the pipeline if the header or HTML contains noindex.
Put the WordPress Search engine visibility box on your launch checklist, and check it again after every migration from a staging copy.
Schedule an external check that runs whether or not anyone remembers, because header rules at the CDN can change without a code deploy.
Get an alert when a noindex or other regression appears
The deploy test catches a noindex introduced by your own code. A scheduled audit catches the rest, including changes made in a CMS or at the CDN without a deploy. LaunchScaler Watch re-scans one website every two weeks for $29/mo, emails a diff on every run showing which checks moved and the evidence behind each one, runs uptime checks on your health endpoint in between, and alerts you when a check drops below the bar. Start watching.
Google drops a page from results when Googlebot next crawls it and sees the noindex, even if other sites link to it. Google does not publish a timeline, because it depends on how often each page is crawled.
03
Why did my site get deindexed after an update?
The most common cause is a deploy or setting change that added a noindex: a staging environment variable reaching production, the WordPress Discourage search engines box being ticked, or a server or CDN rule adding an X-Robots-Tag header. Check the header and the meta tag on your key pages first.
04
How do I get my pages back after removing a noindex?
Remove the noindex at its source, confirm the fix with URL Inspection's live test, request indexing for your most important pages and resubmit your sitemap. Google says crawling can take anywhere from a few days to a few weeks.
05
Does the WordPress Discourage search engines setting block my site?
Since WordPress 5.3 it adds a robots meta tag with noindex,nofollow to every page instead of changing robots.txt. It does not block visitors, and WordPress notes it is up to search engines to honour the request.
SEO monitoring for a small site: uptime every few minutes, indexing directives on every deploy, a technical scan every two weeks, Search Console monthly.