Canonical tags for SEO: setup and the mistakes that break them
A canonical tag is one absolute rel=canonical link in the page's head. The mistakes that make Google ignore it, each with a quick test.
LLaunchScaler·Published ·8 min read
A canonical tag is one line in a page's <head>, <link rel="canonical" href="https://example.com/page">, that tells Google which URL to index when several URLs show the same content. It works when there is exactly one, it uses an absolute URL, it sits in a valid <head>, and it points at a live page that your sitemap and internal links also use.
Google treats rel="canonical" as a strong signal that it weighs against your other signals, with no guarantee it will follow it, so a tag that contradicts them gets overruled. The mistakes below are the ones that make Google ignore the tag, each with a test.
What does a correct canonical tag look like?
A correct canonical is a single <link> element in the <head>, with rel="canonical" and an absolute href that includes https:// and the host. On a standalone page it names the page's own clean URL. On a duplicate, it names the page you want indexed in its place.
Google's documentation ranks the ways to declare a canonical by strength: a redirect is a strong signal, rel="canonical" is a strong signal, and sitemap inclusion is a weak one. The signals stack, so the tag works best when your redirects, sitemap and internal links all name the same URL.
You can also send the canonical as an HTTP header, which is how you set one for a PDF: Link: <https://example.com/guide.pdf>; rel="canonical". Google recommends choosing one method, tag or header, rather than both, because two methods can end up naming two different URLs.
Questions, answered
What people ask about this
01
What is a canonical tag?
It is a link element with rel="canonical" and an href, placed in a page's head, that tells search engines which URL you want indexed when several URLs show the same or very similar content. Google treats it as a strong signal, not a command.
02
Should every page have a self-referencing canonical?
In the Next.js App Router, set alternates: { canonical: 'https://example.com/pricing' } in the page's metadata. With metadataBase set in the root layout, a relative value such as /pricing resolves to the full URL.
Should every page have a self-referencing canonical?
Yes. Google recommends adding the canonical link to the canonical page itself. A self-referencing canonical on every standalone page means that when the page is reached as /pricing?utm_source=newsletter or /pricing?ref=x, the tag still names /pricing, and the variants consolidate onto it.
Self-referencing does not mean copying the requested URL into the tag. The tag must name the clean URL, built from your route, not from the incoming request. A template that echoes the current URL, query string included, gives every tracking variant its own "canonical" and defeats the purpose.
Which canonical mistakes make Google ignore the tag?
Google overrides or ignores a canonical when it is ambiguous, unreachable or contradicted. The usual causes are a relative URL, a target that redirects or returns an error, more than one tag, a tag outside the <head>, and a template that points every page at the homepage. Each has a quick test.
Mistake
What Google does
How to test
Relative or scheme-less URL, such as href="example.com/page"
Reads it as a path and resolves it to https://example.com/example.com/page; Google's 2013 post says it may ignore the tag.
View source and check the href starts with https://.
Canonical points to a redirect
Two signals disagree: the tag names A, the redirect names B, and Google picks one itself.
curl -sI the canonical URL; it must return 200, not 301 or 308.
Canonical points to a 404, a soft 404 or a noindexed page
Google's guidance is that the target must exist and must not carry noindex.
curl -sI for the status, then check the target has no noindex.
Two or more canonical tags
Google has said it will likely ignore all of them.
Count the tags in the rendered HTML (command below).
Canonical in the <body>
Disregarded; only the <head> counts.
Check where it sits in the rendered DOM (console test further down).
Homepage canonical on every page
Every page asks to be treated as a duplicate of /, so Google can fold inner pages into the homepage and drop them.
Fetch three inner URLs and compare each tag to its own URL.
Canonical from page 2 of a series to page 1
Pages 2 and later are not duplicates of page 1, so their content can fail to get indexed.
Check paginated URLs name themselves.
A few commands cover most of these. Count the canonical tags in the served HTML:
One line of output is right. Zero means no tag, and two or more means Google will likely ignore them all. Then check the target answers 200:
curl -sI https://example.com/pricing | head -n 1
Duplicate tags come from two layers each adding one, for example an SEO plugin plus a theme, or a framework's metadata plus a hand-written <link> in a layout. Google's 2013 post on these mistakes names SEO plugins that insert a default canonical as a frequent source, and tells you to check the entire <head>, because the two tags can be far apart.
The homepage-on-every-page mistake is easy to make in frameworks that inherit metadata. In Next.js, metadata from a layout is shallowly merged into every page below it, so an alternates.canonical set in the root layout is inherited by every page that does not set its own. Set the canonical per page, or generate it from each page's route.
Why does invalid HTML in the head drop the canonical?
Google stops reading the <head> at the first element that is not allowed there. Its documentation says that once it detects an invalid element, "it assumes the end of the <head> element and stops reading any further elements." A canonical, robots meta tag or hreflang link that comes after that element is ignored.
The HTML standard allows only these elements in the <head>: title, meta, link, script, style, base, noscript and template. Google names iframe and img as common offenders. They tend to arrive from snippets pasted into the head: a tracking pixel's <img>, a tag manager's <iframe>, or a <div> a widget injects for its own use.
Browsers follow the same rule, because the HTML parsing algorithm in the WHATWG standard closes the head at an element that does not belong there, opens the body, and puts every later tag in the body. That gives you a reliable test. Open the page, open DevTools, and run this in the console:
A result of 1 or more means the canonical was pushed out of the head and into the body. Then look in the Elements panel for the first element in the <head> that is not on the allowed list; everything from there down was moved. The fix is to move the canonical, robots and hreflang tags to the top of the head, directly after <meta charset> and <title>, and move the offending element into the <body>. Google also recommends placing any element you cannot avoid after the ones you want it to see.
Do the canonical and the sitemap have to match?
Yes. Google's documentation tells you not to specify one URL as canonical in the sitemap and a different URL in rel="canonical" for the same page. List only canonical URLs in the sitemap, exactly as the tag writes them: same protocol, same host, same trailing slash.
Mismatches are usually tiny: the sitemap lists https://www.example.com/pricing/ while the tag says https://example.com/pricing. To Google those are two different URLs sending two different preferences. The same rule extends to internal links: Google recommends linking to the canonical URL rather than a duplicate, so your navigation, footer and in-body links should use the form the tag names. The trailing slash guide shows how to pick one form and redirect the other so the tag, sitemap and links cannot drift apart.
If your site has language versions, each version's canonical should name a page in the same language. The hreflang tags guide covers how canonicals and hreflang annotations fit together without cancelling each other out.
How do you know which canonical Google chose?
Inspect the URL in Search Console. The Page indexing section shows the User-declared canonical (what your tag asks for) and the Google-selected canonical (what Google indexed). If they differ, Google overruled your tag, and the page is usually listed in the Page indexing report under one of the duplicate statuses.
Page indexing status
What it means
Guide
Alternate page with proper canonical tag
Your tag pointed elsewhere and Google agreed. Nothing to fix.
Google's troubleshooting guide lists the causes it sees behind an unexpected choice: CMS or plugin canonicals pointing at the wrong URL, hosting misconfigurations, missing hreflang on localized copies, and pages that are simply too similar. It also warns that after you fix the content, Google might hold pages in a duplicate cluster for up to two weeks. Canonical choice is only visible in the indexed result; the live test in URL Inspection cannot predict it.
What should the canonical never be used for?
Don't use a canonical to hide a page, to point a category page at one featured article, or to fold pagination into page one. Google's own guidance draws those lines, and each one removes pages you probably want in results.
A canonical is not a noindex. To keep a page out of Google, use noindex. Google also advises against using noindex to choose a canonical within a site, because it blocks the page from Search entirely.
A category or landing page should not canonicalize to the article it features that day. Google's 2013 post explains that the category page would then stop appearing in results.
Paginated pages each keep their own canonical. Google's pagination guidance says not to use the first page of a series as the canonical for the rest.
Don't use robots.txt for canonicalization. Google says a disallowed URL can still be indexed without its content.
Check the canonical on your live pages
A canonical problem is invisible in the browser and usually comes from a template, so it affects every page built from it. Run the free scan on LaunchScaler with just your address and no account: it runs 156 checks across 6 of its 7 categories at no cost, and its search checks flag a page with no self-referencing canonical or one pointing at a different URL, a canonical whose target redirects, returns 404 or is noindexed, invalid HTML in the <head> that drops the canonical and robots tags below it, and a sitemap that lists non-canonical URLs.
Yes. Google recommends adding a rel="canonical" link on the canonical page itself, so a standalone page names its own clean URL. That protects it when the same page is reached with tracking parameters or other URL variants.
03
Can a canonical tag use a relative URL?
Google supports relative URLs but recommends absolute ones. A path written without the protocol, such as example.com/page, is read as a relative path and resolves to a URL that does not exist.
04
What happens if a page has two canonical tags?
Google has said that when a page specifies more than one rel=canonical, it will likely ignore all of them. It then picks a canonical on its own.
05
Does the canonical tag work in the body?
No. Google only accepts rel=canonical in the head. An invalid element such as an img or iframe in the head makes Google assume the head has ended, so a canonical placed after it is ignored as well.
Google's crawl budget guide is for sites with a million pages, or 10,000 changing daily. What a small site should fix instead, and how to read Crawl Stats.
Crawled - currently not indexed means Google fetched the page and chose not to index it. Triage by URL pattern, then merge, noindex or improve each page.
Discovered - currently not indexed means Google knows the URL but postponed the crawl. Check server capacity in Crawl Stats, then raise demand with links.