Orphan pages: find the pages only your sitemap links to
An orphan page is in your sitemap but no internal link reaches it. Find orphans by diffing the sitemap against a crawl, then link or remove each one.
LLaunchScaler·Published ·7 min read
An orphan page is a URL that is in your sitemap but that no internal link reaches: start at the homepage, follow every link, and you never arrive at it. To find orphans, list every URL in the sitemap, crawl the site from the homepage following internal links only, and compare the two lists. Then fix each orphan by linking it from a relevant page, or remove it from the sitemap if it shouldn't exist.
Orphans appear on sites that publish often, change navigation, or generate landing pages. Nothing breaks, so they go unnoticed until someone asks why a page never ranks.
What is an orphan page?
An orphan page is a live URL on your site with no internal links pointing to it. Crawlers that follow links from your homepage can't reach it, so the only ways a search engine learns about it are your XML sitemap, a link from another site, or its memory of a URL it saw before. Many orphans survive only in the sitemap.
They tend to appear in predictable ways:
A navigation or footer redesign drops a link that was the page's only way in.
A blog listing paginates, and older posts are left linked only from deep archive pages or from nowhere.
Landing pages built for ads or email campaigns were never linked from the site on purpose.
A CMS keeps publishing tag, author or attachment pages into the sitemap that nothing links to.
A page was retired from the menu but never removed or redirected.
The last two kinds are orphans you want gone. The first three are pages you want found.
Why do orphan pages hurt SEO?
Google finds most pages by following links, and it uses those links to judge what a page is about. Its link guidance opens with it: "Google uses links as a signal when determining the relevancy of pages and to find new pages to crawl." An orphan misses both. The sitemap can make the URL known, but no link tells Google it matters or what it covers.
Google's own words set the bar: "Every page you care about should have a link from at least one other page on your site." Its sitemap guide makes the same point from the other side, saying a sitemap helps search engines discover URLs "but it doesn't guarantee that all the items in your sitemap will be crawled and indexed," and that a site whose important pages can all be reached "by following links starting from the home page" may not need one.
Questions, answered
What people ask about this
01
What is an orphan page in SEO?
An orphan page is a URL on your site that no other page links to, so a crawler following internal links from the homepage never reaches it. It is usually still listed in the XML sitemap, which is often the only way search engines know it exists.
In Search Console, orphans can surface in the Page indexing report as "Discovered - currently not indexed" (Google knows the URL but hasn't crawled it) or "Crawled - currently not indexed" (it looked and passed). The discovered currently not indexed guide and the crawled currently not indexed guide cover both statuses.
Pages linked only from noindexed pages are close to orphans as well. Google's John Mueller said in a 2017 Search Central office-hours video that a page kept on noindex long enough is removed completely and its links are no longer followed, so a post reachable only from a noindexed tag archive gradually loses its only path. The noindex, follow guide covers that case.
Bing, whose index Microsoft Copilot answers from, asks for the same thing. Its webmaster guidelines say to ensure each important URL is reachable through crawlable internal links, using standard <a href> links with relevant anchor text.
How do you find orphan pages?
Compare two lists: every URL in your XML sitemap, and every URL a crawler reaches by following internal links from the homepage. URLs in the first list but not the second are orphans. You can do it with a desktop crawler that compares them for you, or with a crawl export and two text files.
With a crawler that reports orphans
Screaming Frog's SEO Spider, free for crawls up to 500 URLs, has an Orphan URLs filter that does the comparison. It defines orphans as "URLs that are only in an XML Sitemap, but were not discovered during the crawl."
In the SEO Spider, enable Config > Spider > Crawl > Crawl Linked XML Sitemaps.
Crawl the site from the homepage.
Run Crawl Analysis when the crawl finishes.
Open the Sitemaps tab and choose the Orphan URLs filter, then Export.
It also has Orphan URLs filters for Google Search Console and Google Analytics data, which catch orphans that aren't in the sitemap but are getting impressions or visits.
With two lists and a diff
Without a crawler that compares for you, build the two lists yourself:
# 1. Every URL in the sitemap (expand a sitemap index into its child sitemaps first)
curl -s https://example.com/sitemap.xml \
| grep -o '<loc>[^<]*</loc>' \
| sed -e 's/<loc>//' -e 's/<\/loc>//' \
| sort -u > sitemap.txt
# 2. Every internal HTML URL your crawler reached, one per line, exported from its results
sort -u crawl-export.txt > crawled.txt
# 3. URLs in the sitemap that the crawl never reached
comm -23 sitemap.txt crawled.txt > orphans.txt
Normalise both lists the same way before comparing: the same protocol, the same host, and the same trailing slash convention. Otherwise https://example.com/pricing and https://example.com/pricing/ show up as two different URLs, and a page that is linked gets reported as an orphan.
Make sure the crawl follows only links a search engine can follow: plain <a href> elements. Google says it generally can only crawl a link if it is an <a> element with an href attribute, so a page reachable only through a JavaScript click handler or a form counts as orphaned even though visitors can get to it. If your navigation is rendered by JavaScript, crawl with rendering turned on and compare that list too.
How do you fix an orphan page?
Decide for each orphan whether it should exist. If it should rank, link to it from a relevant page that is already indexed, using descriptive anchor text. If it shouldn't, take it out of the sitemap, then redirect it, noindex it or delete it. Screaming Frog's own fix line says the same: link the ones that should rank, and remove the rest from the sitemap.
Orphan type
Keep?
Fix
A product, feature or pricing page
Yes
Link it from the main navigation, the homepage, or a relevant hub.
An older blog post
Yes
Link it from two or three related posts and from the category or pillar page for its topic.
A campaign landing page
Your call
If it shouldn't rank, remove it from the sitemap and add noindex; if it should, link it.
A thin tag, author or attachment page
No
Remove it from the sitemap; noindex it or turn the page type off.
A retired page
No
301 it to its replacement, or return 404 or 410, and remove it from the sitemap.
Where you place the link matters. Put it in the body of a page that covers a related topic, in a sentence that explains what the reader will find. Google's link guidance asks for anchor text that is "descriptive, reasonably concise, and relevant," and warns against generic anchors such as "click here" or "read more." A footer link to every orphan technically ends the orphan status but tells Google little about each page.
The internal linking guide for a new blog covers building those links into your publishing routine, so new posts aren't orphaned the day they go live.
How many internal links should a page have?
There is no target number. Google says "there's no magical ideal number of links a given page should contain," and adds that if you think it's too much, it probably is. For a sense of scale, the 2025 Web Almanac found that the median desktop page carries 43 links to its own site and the median mobile page 39, rising to 174 and 161 at the 90th percentile.
Those figures count every internal link on a page, navigation and footer included. What matters for an orphan is the step from zero to one or more contextual links from pages that are themselves indexed and linked. Once every page you care about has at least one, check the pages that matter most and give them links from several related pages.
How do you stop orphans from coming back?
Make internal links part of publishing, and recheck after any change to navigation. Most orphans are created by a release or a redesign, not by a single post.
When you publish, add a link to the new page from at least one existing, related page in the same session.
After any navigation, footer or template change, rerun the sitemap-versus-crawl comparison.
Generate the sitemap from the same source as your routes, so retired pages drop out of it automatically.
Check the Page indexing report monthly for new URLs under Discovered or Crawled but not indexed, and test them for links.
Run the free scan on LaunchScaler with your address and no account. It runs 156 checks across 6 of its 7 categories at no cost, including a page that is under-linked internally, internal links that use generic anchor text, broken internal links and dead sitemap URLs, and a sitemap that lists redirecting, noindexed or non-canonical URLs. It checks the page you give it and your sitemap, so use a crawler's orphan report for the full site-wide list.
They are weak. Google can discover an orphan through the sitemap, but it uses links as a signal of relevance and to find pages, and its guidance says every page you care about should have a link from at least one other page on your site.
03
How do I find orphan pages?
Export every URL in your XML sitemap, crawl the site from the homepage by following internal links only, and compare the two lists. URLs in the sitemap that the crawl never reached are your orphans.
04
Can Google find orphan pages?
Yes, through a sitemap, an external link, or a URL it saw before. Finding a page is not the same as treating it as important, though, and an orphan gets none of the internal links Google uses to understand a page.
05
How many internal links should a page have?
Google says there is no magical ideal number. For scale, the 2025 Web Almanac found the median desktop page has 43 links to its own site, and the median mobile page 39.
Page with redirect means the URL redirects, so Google indexes its target instead. Ignore it unless your sitemap or internal links still use the old URL.
Redirect error means Googlebot hit a loop, a chain that was too long, an overlong URL or a bad URL. Trace every hop with curl, then redirect in one step.
Request indexing lives in URL Inspection and has an unpublished daily quota. Repeat requests don't help; a rejection means the live test found a blocker.