Why are some of my pages not indexed by Google when others are?
Inspect one missing page in URL Inspection. Its result names the branch: unknown, blocked, crawled or discovered but not indexed, or a duplicate.
LLaunchScaler·Published ·9 min read
Some pages get indexed and others don't because Google decides page by page, and each missing page falls into one of five branches: Google has never found it, something blocks it, Google crawled it and passed, Google found it but hasn't crawled it, or Google indexed a different URL in its place. Inspect one missing page in Search Console's URL Inspection tool and the result tells you which branch you are in.
Work through one page first, fix its cause, then check whether the same cause explains the others. Below is the tree, with the check and the fix for each branch.
Step 1: What does URL Inspection say about one missing page?
Open Search Console, paste the full URL of one page you expected to be indexed into the inspection bar at the top, and press Enter. The verdict and the "Page is not indexed:" reason under Page indexing put the page in a branch. Pick a page that matters, not a random one: a product, pricing or key article page.
Map the result to a branch:
What URL Inspection shows
Branch
Go to
Page is not indexed: URL is unknown to Google
Discovery
Branch A
Crawl allowed? No, or Indexing allowed? No, or Page fetch: Failed
Blocked
Branch B
Excluded by 'noindex' tag, Not found (404), Server error (5xx), Soft 404, Blocked due to access forbidden (403)
Blocked
Questions, answered
What people ask about this
01
Why are some of my pages indexed and others not?
Google decides page by page. A missing page is either unknown to Google, blocked from crawling or indexing, crawled or discovered but not yet chosen for the index, or treated as a duplicate of another URL. URL Inspection tells you which of these applies to a given page.
02
How do I find out why a specific page is not indexed?
Duplicate without user-selected canonical, or Duplicate, Google chose different canonical than user
Canonicalised elsewhere
Branch D
Page with redirect, or Alternate page with proper canonical tag
Working as intended
Usually nothing to fix
The default view describes Google's last crawl, so check the Last crawl date. If you changed the page after that date, click Test live URL to see its current state. The URL is not on Google guide reads every line of the panel in detail.
If you can't use Search Console yet, Google's help suggests a site: search for the exact URL, such as site:example.com/pricing. It tells you whether the page is indexed, but not why it isn't.
Branch A: Why is the page unknown to Google?
"URL is unknown to Google" means Google has never seen the URL. It is a discovery problem: no crawled page links to it, no submitted sitemap lists it, and no other site links to it. Google's help puts the rule plainly: to learn about a page, you must submit a sitemap or a crawl request, or Google must find a link to it somewhere.
Check, in order:
Is the page linked from any page Google has indexed? Search your site's HTML for the path. A page reachable only through a search box, a filter or a JavaScript onclick is not linked in a way Google can follow; it needs a plain <a href> link.
Is it in your XML sitemap, and is that sitemap submitted in Search Console or named in robots.txt?
Is the URL in the sitemap exactly the URL you inspected, with the same protocol, host and trailing slash?
Fix it by linking the page from a relevant indexed page with descriptive anchor text, adding it to the sitemap, and then requesting indexing in URL Inspection. Google's link guidance says every page you care about should have a link from at least one other page on your site. Pages that are only in the sitemap and linked from nowhere are orphans; the orphan pages guide shows how to find all of them at once.
Branch B: What blocks a page from being indexed?
A blocked page is stopped by a rule or a response: a robots.txt Disallow, a noindex in the meta tag or X-Robots-Tag header, or an HTTP status that isn't 200. URL Inspection names which. "Crawl allowed? No" points at robots.txt, "Indexing allowed? No" at a noindex, and "Page fetch: Failed" at a 4xx or 5xx response.
Block
How it shows
Fix
robots.txt Disallow
Crawl allowed? No
Remove or narrow the matching rule in the robots.txt of that host.
noindex
Indexing allowed? No, naming the meta tag or X-Robots-Tag header
Remove it from the template, CMS setting or server config that adds it.
404 or 410
Page fetch: Failed, Not found (404)
Restore the page, or 301 it to its replacement.
401 or 403
Blocked due to unauthorized request (401) or access forbidden (403)
Remove the login, or allow verified Googlebot through the firewall. Googlebot never sends credentials.
5xx
Server error (5xx)
Check host status in the Crawl Stats report and the server logs for the crawl date.
Soft 404
Page returns 200 but looks empty or "not found"
Return a real 404 for missing pages; fix rendering if the page is real.
Two combinations mislead. When robots.txt blocks a page, "Indexing allowed?" always shows Yes, because Google can't see any noindex behind the block. And 4xx responses for URLs Google already indexed get them removed: Google's status code documentation says the indexing pipeline removes URLs that return 4xx, and persistent 5xx errors lead to removal too.
A block that hits some pages but not others lives in whatever those pages share: a template, a path rule or a firewall rule. Typical examples are a noindex added to one page type, a Disallow: /blog/tag without a trailing slash that also catches /blog/tag-manager-guide, or a firewall rule scoped to one path.
Branch C: Why is a page crawled or discovered but not indexed?
These two statuses mean nothing blocks the page; Google chose not to index it yet. "Crawled - currently not indexed" means Google fetched the page and passed for now. "Discovered - currently not indexed" means Google knows the URL but hasn't crawled it, which Google's help attributes to a crawl that "was expected to overload the site."
For Crawled - currently not indexed, Google's help says the page "may or may not be indexed in the future" and that there is "no need to resubmit this URL." Resubmitting does not change the outcome, but improving the page can:
Compare it with the pages that are indexed. If it repeats a template with only a few words changed, or covers the same topic as another page on your site, Google has little reason to index both.
Measure it against Google's own questions for helpful content, such as whether it provides "original information, reporting, research, or analysis" and "substantial value when compared to other pages in search results."
Link to it from indexed pages that cover related topics, so Google sees it as part of the site rather than a leftover.
For Discovered - currently not indexed, look at demand and capacity. Open the Crawl Stats report under Settings and check host status and average response time. Google's crawl budget guide says that when a site slows down or returns server errors, Google crawls less. Then check that the page is linked from pages Google crawls often, not only listed in a sitemap.
If URL Inspection shows a Google-selected canonical that differs from the URL you inspected, the page is not missing: Google indexed another URL for the same content. That URL gets the rankings, and yours is treated as its duplicate. Inspect the Google-selected URL to confirm it is indexed.
Compare three things in your browser: the page you inspected, its User-declared canonical and the Google-selected canonical. Then decide:
If Google's choice is the URL you want in results (a clean URL over a parameter variant, https over http), there is nothing to fix. Point internal links and the sitemap at it.
If Google picked the wrong URL, make every signal agree on the right one: a rel="canonical" in the <head>, the sitemap, internal links and, for retired duplicates, a 301.
If the pages are not really duplicates, make them differ substantially. Google's help says a duplicate must be similar to its canonical, so a page Google clusters with another is one it sees as too similar.
Google's canonical troubleshooting guide adds that after you fix content, Google "might hold pages in a duplicate cluster for up to two weeks."
Step 2: Does the same cause repeat across the site?
Once one page is fixed, check whether its cause explains the rest. Open the Page indexing report, find the row that matches the reason you just fixed, and read its example URLs. If they share a template, a folder or a URL pattern, fix the pattern once, then click Validate fix on that row.
Look for these patterns:
A whole folder in one row, such as every /blog/ post under "Crawled - currently not indexed", which points at the template or internal linking rather than individual posts.
A drop in indexed pages with no matching rise in errors. Google's help says this can mean you are blocking existing pages with robots.txt, noindex or a required login.
More non-indexed than indexed pages, which Google's help ties to robots.txt blocking large sections, or to many duplicate parameter URLs (?sort=price, ?color=green).
An error spike after a release, which Google's help says can come from a template change or a sitemap that lists blocked URLs.
The report reflects Google's last crawl of each URL, so it lags behind your fixes, and Google's help says a site with fewer than 500 pages probably doesn't need it at all. The page indexing report guide explains every row and how Validate fix works.
Which missing pages are fine to leave out?
Not every URL should be indexed. Google's help says to expect only the canonical version of each important page in the index, and calls duplicate or alternate statuses "usually a good thing." Leave out redirected URLs, filtered and sorted list views, login and account pages, and anything you noindexed on purpose.
Focus on pages a searcher should land on: the homepage, product and pricing pages, features, and articles you want found. If those are indexed and the rest are missing for expected reasons, the site is working.
Check the blocking causes across every page
URL Inspection answers one URL at a time, on a property you own. To check the causes behind Branch B and Branch D across your live pages, run the free scan on LaunchScaler with just your address and no account. It runs 156 checks across 6 of its 7 categories at no cost, including robots.txt rules that block a page you want ranked, noindex in the meta tag or X-Robots-Tag header, soft 404s, canonicals that point at a redirected, missing or noindexed URL, main content that only appears after JavaScript runs, and a sitemap that lists non-canonical URLs.
Paste its full URL into the URL Inspection bar in Search Console. The Page indexing section names the reason after 'Page is not indexed:', and the Crawl allowed?, Page fetch and Indexing allowed? lines show whether a block is involved.
03
Why does Google crawl a page but not index it?
'Crawled - currently not indexed' means Google fetched the page and chose not to index it for now. Google does not name the cause, so check whether the page is thin, repeats another page, or is barely linked. Google says there is no need to resubmit the URL.
04
Should every page on my site be indexed?
No. Google's help says not to expect 100% coverage: only the canonical version of each important page should be indexed, and duplicates, filtered lists and utility pages are normally left out.
05
How long does Google take to index new pages?
Google's help says it can take a week or so for Google to start crawling and indexing a new page or site, and that indexing is never instant, even after a crawl request.
Check WordPress in this order: the Discourage search engines box, SEO plugin noindex settings, one sitemap, coming-soon mode, firewalls, then cached pages.
Neither www nor non-www ranks better. Pick one host, 301 the other in one hop, and align DNS, HSTS and Search Console so only one copy of your site exists.