noindex, follow: does Google follow links on a noindexed page?
At first, yes. Google says a page left on noindex long enough is dropped and its links stop being followed, so noindex, follow ends up like nofollow.
LLaunchScaler·Published ·8 min read
Google follows the links on a noindex, follow page at first, but not forever. Google's John Mueller has said that once a page has carried noindex for long enough, Google removes it completely and stops following its links, so in the long run noindex, follow works like noindex, nofollow.
That changes how you should use the tag. It is still the right choice for thin archive and search pages you want out of results. It is the wrong way to keep important pages discoverable, because a hub page on noindex eventually stops passing crawlers to anything it links.
What does noindex, follow mean?
<meta name="robots" content="noindex, follow"> asks search engines to keep the page out of their results while still crawling the links on it. The same pair can be sent as an HTTP header, X-Robots-Tag: noindex, follow. For Google, follow adds nothing: following links is the default whenever nofollow is absent.
Google's robots meta tag specification lists noindex, nofollow and none as rules, and defines nofollow as "Do not follow the links on this page. If you don't specify this rule, Google may use the links on the page to discover those linked pages." There is no separate follow rule in that list, because it is the default behavior. So these three lines behave the same for Google:
It asks search engines not to show the page in results while still crawling the links on it. For Google, follow is already the default, so the tag behaves the same as a plain noindex.
Writing follow out does no harm and documents your intent for the next developer. It just does not buy any extra link following, and it does not stop the long-term behavior described next.
Does Google keep following links on a noindexed page?
In the short term, yes. Over time, no. Mueller explained in a Google Search Central office-hours video in December 2017 that Google first keeps a noindexed page in its index without showing it, and can follow its links. If the noindex stays, Google eventually removes the page completely, and then its links are not followed either.
His words, from the transcript of that video (from 54:51):
"But if we see the noindex there for longer than we think this this page really doesn't want to be used in search so we will remove it completely. And then we won't follow the links anyway. So in noindex and follow is essentially kind of the same as a noindex, nofollow. There's no really big difference there in the long run."
When asked afterwards how long that takes, his answer was "It depends." Google has not published a time frame, and crawl frequency plays a part: a page Google rarely visits is seen with its noindex less often. Google's own noindex documentation adds that, depending on a page's importance, "it may take months for Googlebot to revisit a page."
The practical reading is simple. Treat a permanently noindexed page as a dead end for crawling. Any page that depends on it for discovery needs another path from a page that is indexed.
Should a noindexed hub page pass your internal links?
No. Do not rely on a noindexed page to lead Google to pages you want indexed. Link every important page from at least one page that is indexed and stays indexed, such as the homepage, a category page, a pillar article or the page it most relates to. Google's link guidance states it plainly: every page you care about should have a link from at least one other page on your site.
The pattern that breaks goes like this:
A blog template noindexes its tag archives and paginated listing pages to avoid thin pages in results.
Older posts drop off the first listing page and are now linked only from /blog/page/4/ and from tag archives.
Months later, Google has dropped those noindexed pages and stopped following their links.
The older posts now have no crawlable path from an indexed page. Google still knows their URLs from the sitemap, but nothing it indexes links to them any more.
The fix is structural. Link older posts from related posts in the body, from the pillar or category page that covers the topic, and from a curated "start here" or archive page that is indexed. The internal linking guide for a new blog gives a linking structure that holds up as the post count grows, and the orphan pages guide shows how to find posts that only your sitemap still points to.
Where does noindex, follow still fit?
It fits pages you want out of results for good and that are not the only route to anything important. Thin tag archives, internal search results, filtered or sorted list views, and utility pages such as login or thank-you pages are the usual candidates. On those pages, losing link following later costs you nothing.
Page type
noindex?
What else it needs
Tag archives with one or two posts each
Yes, if they are thin
Every tagged post also linked from an indexed category page or related posts.
Internal search results (/search?q=)
Yes, or block crawling in robots.txt instead
Google lists an empty internal search page as a soft 404 example, so either route is better than leaving them indexable.
Sorted or filtered list views (?sort=price)
Yes, or block crawling
Google's pagination guide suggests noindex or a robots.txt block for alternative sort orders.
Paginated pages (/blog/page/2/)
Usually no
Google recommends a unique URL and self-canonical for each page, with a plain link to the next.
Login, thank-you and account pages
Yes
Nothing; they link to nothing that needs discovering.
Category or pillar pages
No
These are the indexed hubs your other pages depend on.
Pagination is where most people reach for noindex, follow and where Google's own guidance points elsewhere. Its pagination documentation asks for a unique URL per page, such as ?page=2, its own canonical URL on each page rather than a canonical to page one, and <a href> links from each page to the next. It also notes that Google no longer uses rel="next" and rel="prev". Keeping paginated pages indexable keeps the crawl path to deep items open.
Why should you never combine noindex with a robots.txt block?
Because the block wins and the noindex is never read. Google's documentation says that for noindex to work, the page "must not be blocked by a robots.txt file." A blocked URL is never fetched, so Google sees neither the noindex nor the links, and the URL can stay in results from external links, without a description.
Teams usually get here by stacking directives to be thorough: Disallow: /tag/ in robots.txt plus noindex, follow in the tag template. The result is the opposite of both intentions. The tag pages can remain indexed as bare URLs, and none of their links are crawled at all, not even in the short window where follow would have applied. Pick one:
To remove pages from results, leave them crawlable and send noindex.
To stop crawling of an endless URL space and you don't care whether a few bare URLs linger, use robots.txt alone.
If a noindex page has already been blocked, remove the Disallow first, let Google recrawl and drop the pages, and only then consider blocking again. The X-Robots-Tag vs robots.txt comparison covers which directive to use for each kind of page and file.
How do you set noindex, follow on a page?
Put the rule in the page's initial HTML or in the response header, never only in JavaScript that runs later. The meta tag goes in the <head>, the header works for any file type, and most frameworks have a metadata setting that writes the tag for you.
The three common forms:
<!-- In the <head> of the HTML response -->
<meta name="robots" content="noindex, follow">
# As a response header, for example in an Nginx location block
add_header X-Robots-Tag "noindex, follow" always;
In the Next.js App Router, set it in the page's metadata: export const metadata = { robots: { index: false, follow: true } } renders <meta name="robots" content="noindex, follow"> in the head.
Two rules from Google's documentation apply whichever form you use. First, a noindex in the original HTML can stop Google rendering the page at all: its JavaScript guide says that when Google encounters the noindex tag it may skip rendering and JavaScript execution, so a script that removes the tag later may not work. Second, when the meta tag and the header disagree, Google applies the more restrictive rule, so a leftover X-Robots-Tag: noindex, nofollow from a server config overrides a template that says noindex, follow.
If you only want Google not to associate your site with one destination or crawl it from your page, you don't need a page-level rule at all. Put rel="nofollow" (or rel="sponsored" for paid links) on that single <a> element and leave the rest of the page alone.
How do you check what a page actually sends?
Check both places a noindex can live, because the header version never shows in the page source. Read the meta tag in the HTML and the X-Robots-Tag header in the response, then confirm what Google saw in URL Inspection, where "Indexing allowed?" names the source of any noindex it found.
If a page you want indexed shows up under "Excluded by 'noindex' tag" in the Page indexing report, the excluded by noindex tag guide traces it to the template, plugin or header rule that adds it.
Check every page's noindex and robots rules at once
Checking templates one URL at a time misses the page you forgot. Run the free scan on LaunchScaler with just your address and no account: it runs 156 checks across 6 of its 7 categories, including noindex in both the meta robots tag and the X-Robots-Tag header on pages meant to rank, a noindex sitting behind a robots.txt block that Google can never read, and pages that are under-linked internally.
At first. Google's John Mueller said in a 2017 Search Central office-hours video that Google keeps a noindexed page for a while and can follow its links, but once the noindex has been there long enough the page is removed completely and its links are no longer followed.
03
How long before noindex, follow is treated like nofollow?
Google has not given a number. Asked directly, Mueller answered that it depends, and crawl frequency is one of the factors.
04
Should I noindex paginated pages?
Google's pagination guidance does not ask for it. It recommends a unique URL and a self-referencing canonical on every page of the sequence, with plain links from each page to the next, so deep items stay reachable through indexed pages.
05
Can I use noindex, follow together with a robots.txt Disallow?
No. A robots.txt block stops Google fetching the page, so it never sees the noindex or the links. The URL can even stay indexed from external links, without a description.
Most 404s in Search Console are harmless. Fix the ones your sitemap or own pages link to, and 301 URLs with backlinks to their closest equivalent page.
Page with redirect means the URL redirects, so Google indexes its target instead. Ignore it unless your sitemap or internal links still use the old URL.