Search and measurement

Sitemaps and canonical URLs explained

A sitemap lists the pages you want found. A canonical address says which version of a page is the main one. Neither guarantees anything, but mistakes in either can cause real problems.

Published 10 October 2026 · 3 min read

Two files and one tag quietly shape how search engines see your site: the sitemap, the robots file and the canonical tag. Developers talk about them as if everyone knew what they meant. Here is what they are.

The sitemap

An XML sitemap is a file that lists the web addresses you want search engines to find, usually at yoursite/sitemap.xml. It is for machines, not people. Google says a sitemap helps it discover pages, particularly for new or large sites, and that it does not guarantee anything is crawled or indexed. A small site that is well linked together may not need one, though it does no harm.

A good sitemap:

  • lists only addresses you want in search, and only addresses that work;
  • uses the same address form as your site, with the same domain and the same trailing-slash style;
  • leaves out pages that redirect, return an error or are marked noindex;
  • is updated when you add or remove pages;
  • is named in your robots file, and submitted in Search Console.

The canonical address

If the same content can be reached at more than one address, such as with and without www, with and without a trailing slash, or with tracking parameters on the end, search engines treat them as copies and choose one version to show. The canonical address is the one you want chosen.

Google describes several ways to signal it. Redirects are the strongest, because they send visitors and crawlers to one address. A rel="canonical" tag in the page is a strong signal. Listing the address in the sitemap is a weaker one. Use the full address, with the domain, and make each page’s canonical point to its own preferred version.

The robots file

The robots file tells crawlers which parts of a site they may visit. It is not a way to keep a page out of search results. Google says that if a page is blocked in robots.txt, it cannot read a noindex tag on that page. To keep a page out of results, leave it crawlable and mark it noindex.

Common mistakes

Mistake Effect
The sitemap lists test or old addresses Search engines look at pages that do not exist or should not be shown
Pages in the sitemap redirect elsewhere Wasted signals, and confusion about the preferred address
Every page’s canonical points to the home page Google may treat all pages as copies of one
Two versions of the site are live (www and non-www) Duplicates compete with each other
The site is blocked in robots.txt after launch Nothing is crawled
Pages are noindex because the site was a test Pages never appear

How it works on this site

Every indexable page here has a canonical tag that points to its own address on the site’s main domain, which is set in one configuration value so that it can change if a custom domain is added later. The sitemap lists only pages that are meant to be found, and the robots file names it. A check script that I run before each release tests those rules. I mention this so you can see what to ask of any developer, including me.

What to ask your developer

  1. Where is the sitemap, and which addresses does it list?
  2. Does every page have its own canonical address?
  3. Is there one version of the site live, with others redirecting to it?
  4. Is anything blocked or marked noindex, and why?
  5. Have you tested it in Search Console?

If you are planning a move to a new address, follow the redirect rules in planning a redesign without losing search visibility. If your pages are not appearing, begin with why your website is not showing on Google.

Sources

CallWhatsApp