Skip to content
SEO Services

How to Get Canonicalization and Duplicate Content Right

Follow an illustrative retailer project from audit to handover and learn how redirects, canonical tags, sitemaps and links work together to fix duplicate URLs.

Samir Haddad Search & Analytics Lead 27 min read 25 views
How to Get Canonicalization and Duplicate Content Right

Canonicalisation and duplicate content come down to one question: when the same page, or nearly the same page, can be reached at several web addresses, which address should a search engine treat as the real one? You answer that question with redirects, the rel=canonical link element, internal links and sitemaps. When those signals agree, ranking signals gather on one URL and crawlers spend their time on pages that matter. When they disagree, which is the usual state of a site that has grown for a few years, links and relevance get spread across near-identical copies, and the search engine picks a version you may not have chosen.

This guide follows one project from the first brief to the final handover. The project is an illustrative example, a composite built to show how the work is actually done. It is not a client story, and every number in it is invented for teaching. Our retailer is a mid-sized outdoor equipment store with about 2,400 products, a blog, a mix of old and new templates, and the usual pile of accidental duplication. We explain each decision as we go, including the ones that look obvious and the traps we deliberately stepped around.

If you own a site, run marketing for one, or are the developer who will actually ship the fixes, the order of work below matters as much as the individual fixes. Canonical tags written on top of an unsettled host, parameter and redirect setup tend to contradict themselves within a month. Doing the work in sequence avoids that.

The illustrative brief: 2,400 products and far too many URLs

The brief was short. Organic traffic to category pages had been flat for a year while the catalog grew. The marketing lead had noticed that Google sometimes showed a product URL with a tracking parameter attached, and a blog post had been republished on a partner's site that now outranked the original for its main query. The in-house developer asked a practical question: "Should we just add canonical tags everywhere?"

The answer was "yes, eventually, but not first." A canonical tag is one signal among several, and the major search engines treat it as a strong hint rather than a command. Google Search Central's canonicalization documentation is explicit that redirects, rel=canonical, sitemap inclusion and internal linking all feed into canonical selection, and that Google may choose a different URL from the one you declare if the other signals point elsewhere. So the project began by measuring how far apart the signals were.

What we were given at kickoff

  • Read access to Google Search Console and Bing Webmaster Tools for the domain property.
  • Thirty days of raw server access logs, exported from the CDN.
  • Staging access and a developer with about two days a week for the project.
  • A list of the blog posts the company had syndicated to partner publications over the previous two years.

We set three measurable goals. First, reduce the number of crawlable URL variants per real page to as close to one as practical. Second, bring the share of important pages where Google's selected canonical matched the declared canonical close to 100 percent. Third, make sure the original blog posts, not partner copies, were the versions that appeared in search. Rankings and traffic were tracked but deliberately not set as targets, because they depend on far more than canonical hygiene.

2,400real product pages in the illustrative catalog
~61,000distinct URLs found in the initial crawl and logs
11average crawlable variants per product page
34blog posts syndicated to partner sites

Finding where the duplicate URLs came from

The first two weeks were diagnosis only, with no changes to the live site. Duplicate content almost never comes from someone copying text. It comes from the machinery of the site: the same template answering at different addresses. The job at this stage is to list every way an address can vary while the content stays the same, and to count how often each variation actually gets crawled.

Three data sources, compared against each other

We used three views of the site, each of which catches things the others miss.

  1. A full crawl with a desktop crawler set to follow every internal link, read canonical tags and record response codes. The crawl shows what the site itself exposes to a bot that starts at the homepage.
  2. Server logs filtered to verified search engine user agents. Logs show what crawlers actually request, including old URLs that no internal link points to any more but that live on in external links, old sitemaps or the crawler's own memory.
  3. Search Console's page indexing report, especially the buckets labeled "Duplicate without user-selected canonical," "Duplicate, Google chose different canonical than user" and "Alternate page with proper canonical tag." These show the outcome of Google's own canonical selection.

Laying the three side by side gave a clear picture. The crawl found about 19,000 URLs. The logs added another 42,000 that no current internal link reached, most of them parameter combinations and legacy protocol variants. Search Console's duplicate buckets confirmed that Google was already folding many of these together, but was not always picking the URL we would have picked.

The duplication inventory

Grouping the URLs by pattern produced a short list of causes. That is typical: tens of thousands of duplicates usually trace back to five to ten mechanisms, and fixing a mechanism fixes every URL it creates.

Source of duplicationExample pattern (illustrative)Share of variant URLsPlanned fix
Faceted navigation filters/tents/?color=green&size=2p&sort=priceAbout 48%Canonical to clean category, limit crawlable combinations
Tracking and campaign parameters/product/alpine-2p/?utm_source=newsletterAbout 17%Self-referencing canonical to clean URL, strip in internal links
Session identifiers/product/alpine-2p/?sid=8f3a...About 12%Move session state to cookies, redirect legacy URLs
Protocol and host variantshttp://, https://www., https:// without wwwAbout 9%Single host, 301 redirects, HSTS
Trailing slash and case/Tents and /tents/About 7%One convention, 301 redirects
Product variant URLs/product/alpine-2p-greenAbout 5%Case-by-case: consolidate or keep distinct
Print and legacy templates/print/product/alpine-2pAbout 2%Remove template, redirect

Two findings shaped everything after. The site responded with a 200 status on both http and https, and on both the www and bare hosts, so every page existed four times before any parameter was added. And the session identifier had been appended to URLs for visitors with cookies disabled, a leftover from a very old checkout. Bots don't keep cookies, so crawlers were given a new session URL on almost every visit.

One discipline mattered here: we did not add a single canonical tag during diagnosis. Changing signals while you are still measuring them makes the baseline useless, and a baseline is the only way to show later that the work did something. We saved the crawl, the log summary and a Search Console export as the week-two snapshot.

What to do and what to avoid with canonicalisation and duplicate content, side by side
Good practice against the usual mistakes, from the sources listed below.

Choosing one host, one protocol and one slash convention

The first fixes were redirects, not canonical tags. A permanent redirect is the strongest canonicalization signal you have, because the duplicate stops serving content at all. Where a variant has no reason to exist, like an http page on a site that runs on https, redirecting it is better than tagging it.

The host decision

The retailer's backlinks were split roughly evenly between the www and bare hosts, and the brand used both in print. Neither choice had a technical advantage. We picked https with www because the CDN configuration and the cookie scope for the checkout subdomain were already built around it, which meant fewer moving parts. The rule we gave the developer was simple: every request to any other host or protocol combination gets a single 301 hop straight to the https www equivalent, keeping the path and query string.

"Single hop" matters. The legacy setup redirected http to https on the same host, then the bare host to www, which gave two hops for anyone arriving at http on the bare domain. Chains work, but every hop adds latency and a chance of error, and long chains may not be followed to the end. We rewrote the rules so each of the three non-preferred combinations resolves in one step.

Trailing slashes and letter case

The platform generated category URLs with a trailing slash and product URLs without one, but accepted both forms for both and returned 200 either way. We kept the existing convention, since changing it would have meant redirecting every URL on the site for no benefit, and added redirects for the other forms. Uppercase paths, mostly from old print catalogs and email campaigns, were redirected to lowercase.

Once redirects were live we enabled HTTP Strict Transport Security with a short max-age at first, raising it after a week without problems. HSTS makes browsers stop requesting http, which removes a class of mixed-protocol links from new referrals. It is not a search signal in itself, but it keeps the http variants from coming back.

If the site sits behind a CDN or edge layer, as this one did, decide where the redirect logic lives and keep it in one place. Rules split between the origin and the edge are a common source of loops and of variants that redirect on one path but not another. Our guide to CDNs and search crawling covers how edge caching and bot handling interact with this.

One trap is worth naming. The developer's first draft redirected everything to the homepage whenever a legacy URL had no exact match. That turns real, linkable pages into soft 404s and throws away their signals. Every redirect in our map went to the closest equivalent page, and URLs with no equivalent returned a 404 or 410.

Self-referencing canonicals on every indexable template

With the host settled, weeks four and five added a self-referencing canonical to every indexable template: category, product, blog post, brand page, store locator and the static information pages. A self-referencing canonical is a rel=canonical link element in the page head that points to the page's own clean, absolute URL. It looks redundant, but it does real work: when the page is reached with a tracking parameter, a session ID, or through a scraper that copies the HTML, the tag still names the clean original.

Implementation rules we wrote into the ticket

  • Absolute URLs only, including protocol and host: https://www.example.com/product/alpine-2p rather than /product/alpine-2p. Relative canonicals work in principle but break easily when a page is served from a staging host or copied elsewhere.
  • Built from the page's canonical route, not from the request URL. The most common implementation bug is a template that echoes the current address, parameters and all, into the canonical tag. That makes every variant self-canonical and consolidates nothing.
  • One canonical per page, in the head. We found a legacy plugin that injected a second canonical tag into the body on blog posts. When a page carries two conflicting canonicals, search engines are likely to ignore both.
  • The target must return 200 and be indexable. A canonical pointing at a redirecting, 404 or noindexed URL sends contradictory instructions.
  • Rendered and raw HTML must agree. The storefront used client-side JavaScript on some templates. We made sure the canonical was in the server-rendered HTML and that no script rewrote it later.

For non-HTML files, mainly the PDF spec sheets the retailer published for each tent and stove, we used the HTTP Link header to declare a canonical where a PDF duplicated an HTML product page. Most of the spec sheets had unique content and were left self-canonical.

Testing before launch

On staging we crawled with parameters deliberately added to every internal link, then checked that every canonical resolved to a clean URL returning 200 with no noindex. Any row where the canonical target differed from the clean version of the crawled URL was a bug. We found three: the brand template, a search results template that should not have been indexable at all, and paginated blog archives, which pointed every page back to page one.

Decision: Paginated archives and paginated category listings got self-referencing canonicals on each page, not canonicals to page one. Page three of a category lists different products from page one, so it is not a duplicate, and pointing it at page one hides the products that appear only on deeper pages. Google no longer uses rel=next and rel=prev as an indexing signal, so each page in the series has to stand on its own.

Taming faceted navigation and tracking parameters

Faceted navigation accounted for almost half the variant URLs, and it is where canonicalisation and duplicate content decisions stop being mechanical. Some filtered views have real search demand ("green two-person tents") and deserve to be indexable pages. Most combinations ("green, two-person, under four pounds, sorted by price descending, page two") have none and should be consolidated.

Sorting the facets into three classes

We exported every filter the platform offered and put each one into a class, using keyword research on the attribute values and the retailer's own merchandising priorities.

ClassExamplesTreatmentReason
Indexable landing facetsTent capacity, sleeping bag temperature rating, brandClean static path, self-canonical, unique title and intro, in sitemapClear search demand and a distinct set of products
Refinement facetsColor, price band, weight bandCanonical to parent category, links left crawlableUseful to shoppers, little standalone demand
Presentation parametersSort order, items per page, view modeCanonical to unsorted default, links made non-crawlableSame products in a different order, never a search target

Only single-facet selections could become indexable landing pages. Two-facet combinations such as "brand plus capacity" were considered case by case, and only six earned a static page. Anything with three or more facets canonicalized to the nearest indexable parent.

Why canonicals alone were not enough

A canonical tag consolidates signals, but the crawler still has to fetch a URL to read the tag on it. With tens of thousands of filter combinations, the crawler kept rediscovering combinations faster than it could process them. So we also cut the number of crawlable combinations: sort and view controls became form controls that update the page without producing new anchor links, and parameter order was normalized so that ?size=2p&color=green and ?color=green&size=2p could not both exist.

We considered blocking parameter URLs in robots.txt and rejected it for the refinement facets. A URL blocked in robots.txt cannot be crawled, so its canonical tag is never read, and any links pointing at it cannot be consolidated into the parent. Robots.txt controls crawling, not canonicalization. We used it only for one parameter family, internal site search results, which had no external links and no value in the index.

Tracking parameters and session IDs

UTM parameters were handled by the self-referencing canonicals already in place, plus a rule that internal links never carry them. Session IDs were moved out of the URL entirely: the checkout now keeps state in a cookie and, for the tiny share of visitors who block cookies, a server-side fallback that doesn't expose anything in the address. Legacy URLs carrying a sid parameter were 301-redirected to the clean URL with the parameter removed.

Trap avoided: Adding noindex to the refinement facets "to be safe" alongside a canonical to the parent. The two signals contradict each other: one says "this is a copy of that page, merge them," the other says "drop this page." Search engines usually resolve the conflict unpredictably, and over time a long-noindexed page may be crawled less and its links followed less. Pick one intent per URL.

Product variants: when two similar pages are not duplicates

About five percent of variant URLs came from product variants. The platform could publish each color or size of a product as its own URL, and over the years some merchandisers had done so and others had not. This was the least mechanical part of the project, because the right answer depends on how people search and what differs between the pages.

The rule we used

If a variant differs only in a detail the searcher doesn't specify in queries, such as a color on a tent where shoppers search by model and capacity, the variant URL was canonicalized to a single product page with a variant selector. If the variant is a meaningfully different product that people search for by name, such as a three-season versus a four-season version of the same tent line, with different weights, prices and use cases, it kept its own self-canonical page with its own copy.

The test we applied to borderline cases was practical: would a shopper who landed on the wrong variant be annoyed? If a hiker searching for the four-season model landed on the three-season page, yes. If someone searching for the tent by model name landed on the green version rather than the orange one, no, because they can switch color in one click.

Near-duplicate copy

Separate variant pages that were kept distinct needed distinct content, not just distinct URLs. Many had identical descriptions with one adjective changed and identical title tags. We rewrote the copy to lead with what actually differs and fixed the titles at the same time, following the approach in our article on fixing duplicate titles across pages. Where near-identical pages compete for the same query, a canonical tag alone will not make them useful. The pages themselves need to differ in ways that help the reader, which is the core of sound on-page optimization decisions.

We also found around forty discontinued products whose pages were near-empty copies of their replacements. Rather than canonicalizing them to the new models, which would be a mismatch since the products are different, we redirected the ones with a direct successor and removed the rest, using a process similar to the one in our guide on content pruning and refreshing.

Syndicated blog posts and cross-domain canonicals

The retailer had syndicated 34 blog posts, mostly trail guides and gear maintenance articles, to partner publications. In several cases the partner's copy ranked above the original. That is a common result of syndication without agreed signals: the partner site is often larger and better linked, so when the search engine sees two copies of the same article it may treat the partner's as the main one.

What we asked partners to do

The major search engines support cross-domain canonicals for syndicated content, so the preferred fix was to ask each partner to add a rel=canonical on their copy pointing to the original on the retailer's site. Where a partner's CMS couldn't do that, the fallback was a noindex on the partner's copy, which keeps the article off search results while it stays available to the partner's readers. The weakest fallback, used only when a partner would do neither, was a clear, followed link from the syndicated copy back to the original near the top of the article.

Of the 34 posts in our illustrative project, partners added cross-domain canonicals on 21, noindex on 6, and attribution links only on 7. For the future, the retailer's syndication agreement now names the canonical requirement as a condition of republishing and asks for a delay of a week or two between original publication and syndication, so the original is crawled and indexed first. Our practical guide to content syndication and canonicals covers the agreement terms in more detail.

One partner had changed a few headings and trimmed the intro. A cross-domain canonical is still appropriate when the copy is substantially the same article. When a partner rewrites heavily, it is no longer syndication, and a canonical claiming the pages are equivalent would be inaccurate; the right answer then is a credit link and genuinely different content on each side.

We also did not ask partners to take their copies down, even where the partner outranked the original. The syndication brought referral traffic and brand exposure the retailer valued. The goal was to get search engines to credit the original, not to end the partnerships.

Making sitemaps, internal links and canonicals tell the same story

By week seven the canonical tags were correct, but a canonical is only one voice. If the XML sitemap lists URL A, internal navigation links to URL B, and the canonical on both says URL C, the search engine has to weigh three conflicting claims. Bing Webmaster Tools guidance makes the same point as Google's: consistent signals make the preferred URL easy to identify, and mixed signals leave the choice to the engine.

Sitemap cleanup

The old sitemap was generated from the product database and included variant URLs, uppercase paths, http addresses from a hard-coded base URL setting, and several hundred URLs that now redirected. We rebuilt it with one rule: a URL appears in the sitemap only if it returns 200, is indexable, and is its own canonical. Everything else was dropped. The rebuilt sitemap listed about 3,900 URLs against the old one's 11,000-plus.

Internal link cleanup

Internal links are easy to overlook because they are scattered across templates, editorial content and navigation menus. We searched the database and templates for links to non-canonical forms and found them in four places: breadcrumbs that included the active filter parameters, a "recently viewed" widget that carried session IDs, blog posts with hard-coded http links, and email-campaign landing links pasted into the CMS with UTM tags attached. Each was fixed at the source rather than covered with a redirect, because a redirect costs a hop on every crawl and every click.

Other places a URL gets declared

  • Hreflang annotations, if the site has language or regional versions, must reference canonical URLs, and each regional page should be self-canonical rather than canonicalized to another language.
  • Structured data with a url property, such as Product or Article markup, should use the canonical URL.
  • Open Graph tags (og:url) are not a search canonical signal, but keeping them consistent prevents social shares from spreading the wrong URL.
  • Paid media and email templates should link to canonical URLs, with tracking in parameters that the canonical tag already handles.
  1. Weeks 1-2 Crawl, log analysis and Search Console export; duplication inventory and baseline snapshot, no live changes.
  2. Week 3 Host, protocol, slash and case redirects shipped as single-hop 301s; HSTS enabled with a short max-age.
  3. Weeks 4-5 Self-referencing canonicals on all indexable templates, tested on staging with parameterized crawls.
  4. Weeks 5-6 Facet classification, parameter normalization, session IDs moved out of URLs, sort controls made non-crawlable.
  5. Week 6 Product variant review, copy and title rewrites, discontinued pages redirected or removed.
  6. Weeks 6-8 Syndication partners contacted; cross-domain canonicals, noindex or attribution links agreed.
  7. Weeks 7-8 Sitemap rebuilt, internal links fixed at source, structured data and hreflang URLs aligned.
  8. Weeks 9-12 Monitoring of selected canonicals, re-crawl and log comparison, handover of the maintenance routine.

Checking which canonical the search engine actually chose

A canonical you declared and a canonical the search engine selected are two different things, and only the second one affects what appears in search. Checking the gap between them is the final and most important step. It is also where most in-house canonical work stops too early.

URL Inspection, sample by sample

Google Search Console's URL Inspection tool shows both the "User-declared canonical" and the "Google-selected canonical" for an indexed URL. We built a sample of 200 URLs stratified across templates: 60 products, 40 categories, 30 landing facets, 30 blog posts, 20 syndicated posts, and 20 known variant or parameter URLs. We inspected each at week eight and again at week twelve. Bing Webmaster Tools' URL inspection gave a comparable check on a smaller sample.

At week eight, 81 percent of the sample showed matching canonicals. By week twelve that had risen to 94 percent. The remaining mismatches fell into three groups, each with a different cause.

  • Syndicated posts where the partner added only an attribution link. Google still preferred the partner's copy for four of the seven. That is expected: a link is a much weaker signal than a canonical.
  • Two landing facets whose product lists were nearly identical to their parent category, because the attribute was held by almost every product. Google treated them as duplicates of the parent. Google was right. We removed the landing pages and canonicalized them to the parent.
  • Several products where Google selected the old variant URL, which still had strong external links from a gear review site. We kept the redirect in place and asked the reviewing site to update its link; the selection shifted to the canonical over the following weeks.

Reading the page indexing report

At the aggregate level, the page indexing report told the same story. "Duplicate without user-selected canonical" fell sharply once self-referencing canonicals were live. "Alternate page with proper canonical tag" grew, which is good news: those are variant URLs Google found, read and correctly folded into the canonical. The bucket to watch is "Duplicate, Google chose different canonical than user," because each URL in it is a place where your signals are either wrong or outweighed.

Trap avoided: Treating a mismatch as Google "ignoring" the tag and responding by adding more force, such as noindex or robots.txt blocks, on the unwanted version. A mismatch is information. In this project it pointed to a missing partner canonical, two landing pages that really were duplicates, and a strong external link. Each needed a different fix, and none of them was more blocking.

Common misconceptions we corrected along the way

Several beliefs came up during the project that are common enough to address directly, because they lead teams to either overreact or do nothing.

"Duplicate content gets you penalized"

Duplication caused by URL mechanics is a consolidation problem, not a punishment. The search engine picks one version to show and folds the rest into it. The cost is indirect: signals are split, crawling is wasted on copies, and the version shown may not be the one you want. Deliberately scraped or spun content meant to manipulate rankings is a different matter, dealt with under spam policies, and it is not what canonical tags are for. The quality side of that line, meaning whether the page deserves to rank at all, is covered in our pieces on helpful, people-first content.

"Just point everything at the homepage"

We have seen sites where a misconfigured plugin sets every page's canonical to the homepage. The effect is to tell search engines that the entire site is one page. Engines usually recognize this as an error and ignore the tags, but the declared signal is still wrong, and in the worst cases deep pages drop out of the index. A canonical must point to a page with substantially the same content, not to the page you would most like to rank.

"Canonical means removed from search"

A canonical consolidates; it does not remove. If a page must disappear from search because it is outdated, private or legally sensitive, you need noindex, access control or a removal process, as described in our guide to removing content from Google Search. Using a canonical for removal is unreliable because the engine may decide the pages are not equivalent and index the one you wanted gone.

"Once the tags are right, we're done"

New duplication appears whenever someone adds a filter, launches a campaign with new parameters, installs an app that rewrites URLs, or migrates a template. Canonical hygiene is maintenance, not a one-time project, which is why the handover below matters as much as the fixes.

What changed by week twelve, and how the retailer keeps it that way

At the week-twelve review we compared the new crawl, logs and Search Console exports with the week-two baseline. The numbers below are illustrative, as is everything in this project, but the direction and relative sizes are typical of a site whose duplication came mainly from parameters and host variants.

Measure (illustrative)Week 2 baselineWeek 12
Distinct URLs requested by search crawlers in 30 daysAbout 61,000About 9,500
Share of crawler requests going to canonical URLsAbout 22%About 71%
URLs in XML sitemap11,2003,900
Sampled URLs where selected canonical matched declaredNot measured (no declared canonicals on most templates)94%
Syndicated posts where the original is the selected canonical9 of 3427 of 34
Redirect chains of two or more hopsAbout 1,8000 found in crawl

The remaining non-canonical crawl activity was mostly crawlers revisiting redirected legacy URLs, which tapers off over months, and refinement facet URLs that are still crawlable on purpose. We deliberately did not report ranking or revenue changes as outcomes of this project. Consolidation often helps category and product pages, but it happened alongside seasonal demand, merchandising changes and new content, and assigning the credit precisely would be guesswork.

The maintenance routine

  • Monthly: review the "Duplicate, Google chose different canonical than user" bucket in Search Console and inspect any new pattern.
  • Monthly: re-inspect a rotating sample of 30 to 50 URLs across templates and log declared versus selected canonicals.
  • Before any template or platform release: crawl staging with parameterized links and confirm every canonical resolves to a 200, indexable, self-canonical URL.
  • Whenever a new filter or parameter is added: classify it as landing, refinement or presentation before it goes live.
  • Quarterly: compare crawler requests in logs against the sitemap to spot new variant families.
  • For every new syndication deal: confirm the partner's cross-domain canonical within a week of their copy going live.
  • After any host, CDN or HTTPS change: test all protocol and host combinations for single-hop 301s.

The retailer's developer owns the release checks, and the marketing lead owns the monthly Search Console review and partner follow-ups. Splitting ownership that way matches who can actually act on each kind of problem.

When to handle canonicalization in-house and when to bring in help

Much of this project could be done by a competent in-house team with a crawler, access to logs and a developer's time. Self-referencing canonicals, a single host with clean redirects, and a disciplined sitemap are within reach of most teams once they know the rules. The work gets harder, and outside help pays for itself, in three situations that match the brief's own warning signs.

  • Search Console reports the wrong canonical at scale. A handful of mismatches is normal. Hundreds on important templates usually means signals are conflicting somewhere you haven't looked yet, such as a CDN rule, a JavaScript rewrite or an old sitemap still being fetched.
  • Parameters have multiplied URLs faster than you can classify them. Faceted navigation on a large catalog needs decisions about search demand, crawl control and platform capability together, and mistakes are expensive because they apply to thousands of pages at once.
  • Content is syndicated across domains. Cross-domain canonicals depend on partners, contracts and CMS limitations outside your control, and the fallback options each carry trade-offs.

When those apply, a specialist canonicalization and duplicate content service can run the diagnosis, write developer-ready tickets and handle the verification loop. Canonical work also rarely stands alone. It tends to surface thin, near-duplicate and outdated pages that need editorial decisions, so it fits naturally within broader SEO services covering both technical and content work.

Decision: In the illustrative project, the retailer kept the ongoing checks in-house and brought outside help back only for the next platform migration. That is a sensible split for most mid-sized sites: specialists for diagnosis, facet strategy and migrations, and the in-house team for the routine monitoring that catches new duplication early.

Where this comes from

The figures and practices above come from the sources listed.

Working on something like this?

We take on SEO Services work for teams who want it done once, properly. Tell us what you are building and we will tell you honestly whether we are the right studio for it. Start a project.

Where to go next

Spotted something wrong? Report an error on this page. We correct on the page and say what changed.

Frequently asked questions

No. Duplication created by URL parameters, host variants or pagination is treated as a consolidation problem: the search engine picks one version and folds the others into it. The real costs are split ranking signals, wasted crawling and the wrong URL showing in results. Deliberately copied or spun content meant to manipulate rankings falls under spam policies, which is a separate issue.
No. Google treats rel=canonical as a strong hint, not a directive, and weighs it alongside redirects, sitemap inclusion, internal links and external links. If those signals point elsewhere, or the pages are not really equivalent, Google may select a different canonical. You can see its choice in the URL Inspection tool in Search Console.
Generally yes. A self-referencing canonical names the clean URL even when the page is reached with tracking parameters, session IDs or through scraped copies. It should be an absolute URL, appear once in the head, and be generated from the page's own route rather than the address that was requested.
Usually not. Page two or three of a category lists different items from page one, so it is not a duplicate, and pointing it at page one can hide products that only appear deeper in the series. Give each paginated page a self-referencing canonical and make sure the items on it are reachable through normal links.
Robots.txt controls crawling, not canonicalization. If a duplicate URL is blocked, the crawler cannot read its canonical tag, so any signals pointing at that URL cannot be consolidated into the preferred version. Reserve robots.txt for URL families with no value to search at all, such as internal site search results.
Ask the partner to add a cross-domain rel=canonical on their copy pointing to your original. If their platform cannot do that, a noindex on their copy is the next best option, and a clear link back to the original is the weakest fallback. Putting these terms in the syndication agreement and delaying republication until your original is indexed both help.
Consider specialist help when Search Console shows the wrong canonical being selected across many important pages, when faceted navigation or parameters have multiplied URLs faster than you can classify them, or when content is syndicated across domains. These cases involve conflicting signals that are hard to trace and decisions that affect thousands of URLs at once.
All services

The work behind this article, and what it costs.

Samir Haddad

Technical SEO and measurement. Writes about crawling, indexing, Core Web Vitals and the difference between a figure and a guess.

Keep reading