Duplicate content: what counts and what doesn't
There is no duplicate content penalty. That is not reassurance — it is the setup for what does happen, which is consolidation you did not choose.
· 5 min read
There is no penalty, and that is not the good news
Start with the correction, because it is genuinely true and widely misunderstood: there is no duplicate content penalty. Google has said so for years. No sanction is applied to a site for having similar pages, there is nothing to appeal, and the anxiety many small businesses carry about accidentally repeating a paragraph is misplaced.
But people hear that and conclude duplication does not matter, which is the wrong conclusion drawn from a correct fact. What actually happens is not punishment; it is selection. Faced with several pages that say substantially the same thing, a search engine picks one to show and sets the others aside. Nothing is penalised. Something is chosen, and the choosing is done without your input — which means the page representing you for a query may not be the page you would have picked, may not be the one with your booking form, and may change over time without anyone touching the site. Consolidation you did not choose is the accurate description, and it is a real cost with no notification attached.
What actually counts as duplication
The threshold is higher than most people fear. Reusing your standard delivery-terms paragraph across product pages is not duplication in any meaningful sense; neither is a shared footer, a repeated disclaimer, or two articles on related subjects that share vocabulary. Search engines handle these constantly and have to be tolerant of them, because almost every site on the web reuses boilerplate.
What counts is when the substantive content of two URLs is the same or nearly so — when the part of the page that answers the visitor's question is interchangeable. That includes the same page reachable at several addresses, which is the largest and least visible category: with and without a trailing slash, http and https, with and without www, with tracking parameters appended, with query parameters in different orders. A single article can exist at eight addresses without anyone publishing it more than once. It also includes location pages differing only by a place name, product variants differing by a colour, and printer-friendly versions of the same article. The distinction that matters is whether a visitor landing on either would notice a difference. If not, an engine has two candidates for one job.
Where it comes from without anyone publishing twice
The most common sources are technical rather than editorial, which is why the problem persists on sites where nobody has ever copied anything. URL variants are first: unless the site consistently redirects to one canonical form, every page has siblings. Tracking parameters are second — a campaign link appending a source parameter creates a new URL serving identical content, and if those get linked or shared, they become known addresses.
E-commerce platforms generate more. Faceted navigation produces a distinct URL for every combination of filters, most of which show a subset of the same products. Session identifiers in URLs used to be a notorious source and still appear. Pagination can produce pages listing the same items in a different order. Content management systems contribute tag and category archives that reproduce article text in full rather than excerpts, and print stylesheets implemented as separate URLs. Syndication adds the external case: an article republished elsewhere with permission, where a search engine now has two copies on different domains and may select the other one. None of this involves anyone writing anything twice.
Canonical tags, and what they actually do
A canonical link in a page's head names the URL you consider the preferred version of that content. It is the main instrument for this problem and it is routinely misunderstood in two directions. It is a hint, not a command: you are telling a search engine which version you would like indexed, and the engine weighs it against other evidence and may select a different page anyway. In Search Console you can see the canonical you declared alongside the one Google actually selected, and they do not always agree.
The second misunderstanding is what it means. A canonical asserts that these URLs are the same content. It is not a way to point a weak page at a strong one, or to consolidate two genuinely different pages you would rather were one. Pointing a canonical at a page that is not equivalent gives an engine contradictory information and produces behaviour you will struggle to debug later. Two practical notes: every page should carry a self-referencing canonical naming its own preferred URL, which handles the parameter variants automatically; and a canonical pointing at a page that redirects, or is itself canonicalised elsewhere, forms a chain that may simply be ignored.
Choosing between canonical, redirect and noindex
Three tools, three different situations, and picking the wrong one is the usual error. Use a redirect when the duplicate does not need to exist for anybody: an old URL after a restructure, a second version of a page you have consolidated. A permanent redirect removes the duplicate entirely and preserves anyone arriving at the old address. This is the strongest option and the right default when the page has no reason to remain.
Use a canonical when the duplicate genuinely needs to exist for visitors but not for search: a filtered product view a customer relies on, a printer-friendly version, a URL with tracking parameters. The page stays available and you have named your preference. Use noindex when the page must exist for visitors and should not appear in results at all, such as internal search results or a thank-you page — but never combine it with a robots.txt block on the same URL, because a crawler prevented from fetching the page never reads the noindex and the page persists indefinitely. And note that noindex and canonical are alternatives rather than partners: a page telling an engine both to index a different URL and not to index this one is giving mixed instructions.
What you can see yourself, and what you cannot
A surprising amount is free to check. You can request your own pages with and without a trailing slash, with and without www, over http and over https, and see whether each variant redirects to one form or serves content at all — four requests per page and no tooling. You can read the canonical out of any page's source and confirm whether it names that page or another. You can append a nonsense query parameter to a URL and see whether the site serves the same page at the new address without a canonical correcting it, which is the fastest way to demonstrate the parameter problem on your own site. You can sort your page titles and find location or variant pages that differ only by a word.
What you cannot see from your own markup is which URL Google actually selected as canonical, or whether the selection changed. That is visible for free in Search Console for a verified property, in the URL inspection tool and the page indexing report, and nowhere else you have access to. Nor can you determine from your own pages whether another site has copied your content — establishing that needs either a manual search for a distinctive sentence from your page, which is free and works reasonably well, or a paid monitoring service. So the technical half is inspectable today at no cost, the engine's actual decision needs a credential somebody has to connect, and the external half is a manual search or a purchase.
Common questions
Will I be penalised for duplicate content?
No. There is no penalty and nothing to appeal. What happens instead is selection: an engine picks one of your similar pages to represent the content and sets the rest aside, without asking which you would have preferred. The cost is losing that choice, not a sanction.
Is reusing the same paragraph across pages a problem?
Generally not. Boilerplate such as delivery terms, disclaimers and footers appears on almost every site and engines handle it routinely. The concern is when the substantive part of two pages — the part answering the visitor's question — is interchangeable, so a visitor would not notice which one they landed on.
Can a canonical tag point at a page that is not really the same?
It can be written that way and it should not be. A canonical asserts the two URLs hold the same content, so pointing one at a merely related page gives contradictory information and produces indexing behaviour that is hard to diagnose later. If the pages differ, differentiate or redirect instead.
Should I use noindex and a robots.txt block together to be thorough?
No, and this combination is a common trap. Blocking the URL in robots.txt prevents the crawler from fetching the page, so it never reads the noindex telling it to drop the page, which then stays in results indefinitely. Allow the crawl, serve the noindex, and let the page fall out.
Related pages