Joomla XML Sitemap for an Online Store

What the file is for, the tags search engines read and the two they ignore, which pages belong in it, the Joomla article trap, and a ten-minute check at the end.

An XML sitemap is a list of the pages you want search engines to crawl, with the date each one last changed. A crawler normally finds pages by following links. The sitemap hands it the list directly, so a product added this morning does not have to wait until something links to it.

It is a hint and nothing more. Google’s sitemap overview says a sitemap helps search engines discover URLs but does not guarantee that everything in it will be crawled and indexed. The same page says who gains most: large sites, new sites with few external links, and sites with many images. A shop with a growing catalogue on a young domain is all three.

What Joomla ships

Nothing. Joomla has no sitemap generator of its own: there is none in a 6.1 installation, and none in the 5.4 and 6.1 development branches at the time of writing. That leaves three routes. Install a sitemap extension, maintain a file by hand, or let a component serve the sitemap for the content it owns. The rest of this post applies to all three, because the file they have to produce is the same.

The file itself

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/shop/oak-desk</loc>
<lastmod>2026-09-14T08:24:15Z</lastmod>
</url>
<url>
<loc>https://www.example.com/index.php?option=com_content&amp;view=article&amp;id=12</loc>
<lastmod>2026-08-02</lastmod>
</url>
</urlset>

The sitemap protocol requires only url and loc. lastmod is optional and takes a W3C date, with or without the time. The protocol also defines changefreq and priority, and you can leave both out: Google’s sitemap guide says it ignores them, and Bing says the same.

Three rules from the protocol catch people out. The file must be UTF-8. A single file may hold no more than 50,000 URLs and be no larger than 50MB uncompressed. And five characters must be escaped in every value, the ampersand among them. The second entry above shows why that matters on Joomla: a site without search-engine friendly URLs has an & in every address, and each one must be written &amp; or the file does not parse.

Serve the sitemap from the root of the site. Google’s guide says a sitemap affects only the pages at or below its own folder, so one at the root can cover everything.

One index, several files

A sitemap index is a sitemap of sitemaps. You submit the index alone and the crawler fetches the files it names. This one comes from a Solidshop store, with the host name changed and the list shortened:

<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.example.com/index.php?option=com_solidshop&amp;view=sitemap&amp;format=xml&amp;type=products&amp;from=1&amp;to=1000</loc>
<lastmod>2026-09-26T08:24:15Z</lastmod>
</sitemap>
<sitemap>
<loc>https://www.example.com/index.php?option=com_solidshop&amp;view=sitemap&amp;format=xml&amp;type=categories</loc>
<lastmod>2026-09-26T08:24:15Z</lastmod>
</sitemap>
<sitemap>
<loc>https://www.example.com/index.php?option=com_solidshop&amp;view=sitemap&amp;format=xml&amp;type=content.articles&amp;from=1&amp;to=1000</loc>
<lastmod>2024-10-08T07:29:45Z</lastmod>
</sitemap>
</sitemapindex>

The size limit is the official reason to split, and few shops reach it. The better reason is diagnosis. Split by kind of page (products, categories, articles) and a problem shows up against the file it belongs to: a sitemap that cannot be fetched, or one that lists far fewer URLs than you expected, points at one part of the site. Google’s page on large sitemaps adds one constraint: the files an index names must be on the same site as the index.

What belongs in it

Only the canonical address of a page

Google tries to crawl each URL exactly as listed, and asks for the full address including scheme and host. So list the one address you want in the results, and make it match the page’s own <link rel="canonical"> character for character: same scheme, same host, same trailing slash. A sitemap that says http://example.com/ while the page says https://www.example.com/ gives the crawler two answers to one question. Sorted and filtered variants of a listing page stay out.

Only pages a guest can open and a search engine may index

The cart, checkout, account pages, the sign-in form and search results have no place in a sitemap. Neither does anything unpublished, anything outside its publishing dates, or anything that needs a login. A page set to noindex is the common mistake: listing it asks for a crawl and refuses the result in the same breath, and Search Console’s Page indexing report then files it under “URL marked ‘noindex’”.

The Joomla trap: an article no menu item leads to

Joomla builds an article’s URL from the menu. If a menu item points at the article, or at a category above it, the URL runs through that item and is the same wherever the link appears. If none does, Joomla falls back to the menu item of the page doing the linking, or to the home page. The same article then answers at several addresses, and which one a sitemap generator prints depends on where it happened to be running.

There is no correct URL to submit for such an article, so the fix is in the menu, not the sitemap. Give the article’s category a menu item. It can sit in a menu that is displayed nowhere; its job is to give every article beneath it one stable address.

An honest lastmod

lastmod is the one optional tag worth the effort, and only if it is true. Google uses it when it is “consistently and verifiably” accurate. Its 2023 note on the subject describes what happens otherwise: tell it a page changed yesterday when it changed seven years ago, and eventually it stops believing your dates.

The rule is the date the content of the page last changed in a way that matters: the main text, the structured data, the links. It is never the date the sitemap was generated. Bing warns against exactly that. And where there is no honest answer, leave the tag out. Google says that is fine for a home page or a listing that only gathers other pages.

Images and languages

Product images

<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
        xmlns:image="http://www.google.com/schemas/sitemap-image/1.1">
<url>
<loc>https://www.example.com/shop/oak-desk</loc>
<image:image>
<image:loc>https://www.example.com/images/shop/oak-desk.jpg</image:loc>
</image:image>
</url>
</urlset>

An entry can name the images on its page, which helps them into image search. Google’s image sitemap page documents two tags today, image:image and image:loc. The caption, title, licence and location tags of older guides were withdrawn in 2022.

Translations

<url>
<loc>https://www.example.com/en/shop/oak-desk</loc>
<xhtml:link rel="alternate" hreflang="en-gb" href="https://www.example.com/en/shop/oak-desk"/>
<xhtml:link rel="alternate" hreflang="fr-fr" href="https://www.example.com/fr/boutique/bureau-chene"/>
<xhtml:link rel="alternate" hreflang="x-default" href="https://www.example.com/en/shop/oak-desk"/>
</url>

On a multilingual site each language version gets its own entry, and each entry lists every version of the page, itself included. The links must be mutual: if the French page does not point back at the English one, Google ignores the pair. x-default names the version for visitors whose language matches none. The urlset needs the xmlns:xhtml="http://www.w3.org/1999/xhtml" namespace. On Joomla the pairs come from language associations, so an article whose translation was never associated has no alternates to list.

The worked example: a Solidshop store

Solidshop serves the sitemap itself, without a separate sitemap extension. It is on by default, and the index lives at:

https://www.example.com/index.php?option=com_solidshop&view=sitemap&format=xml

Products and product categories have been in it since version 1.2.0. From 1.6.0 the bundled Solidshop - Site Content Sitemap plugin adds the site’s menu items, articles and article categories to the same index. Every rule above is applied for you. Only guest-visible, published pages are listed. Anything set to noindex is left out, and so are menu items for the cart, checkout, account, sign-in and search. Articles that no menu item leads to are skipped, not guessed at. An article’s lastmod is the later of its last edit and the moment it went live, and menu items carry none. Translations are linked through Joomla’s associations. Other extensions can add their own pages to the same index.

Options → Sitemap holds the switches: Include category pages, Include product images, URLs per sitemap file (1,000 by default) and Serve at /sitemap.xml. The last is off by default because another extension may already answer at that address. Decide before you submit, since a URL you have given a search engine has to keep working, and note that a real sitemap.xml file in the site root always wins. The go-live checklist has the full list of what is included and how to exclude a menu or a category.

Tell the search engines

  1. Add a Sitemap: line to robots.txt. Every crawler reads it, no account needed. Joomla robots.txt for an Online Store covers the line and the file around it.
    Sitemap: https://www.example.com/index.php?option=com_solidshop&view=sitemap&format=xml
  2. Submit it in Google Search Console. Paste the URL into the Sitemaps report. The report then shows whether the file was read, when Google last fetched it, and how many URLs it found.
  3. Submit it in Bing Webmaster Tools, under Sitemaps. Bing says it fetches a submitted sitemap at once and then returns at least once a day.

Older guides add a fourth step, a “ping” request to Google after every change. Google announced the end of that address in 2023 and has since switched it off. A script that still calls it does no harm and no good.

Check your work

  1. Fetch the headers. The last response should be 200 with an XML content type. A multilingual site may answer with a redirect first, which crawlers follow.
    curl -sIL "https://www.example.com/index.php?option=com_solidshop&view=sitemap&format=xml" | grep -i -E "^HTTP|^content-type"
  2. Count what the index names, then what one file lists. Compare the second number with the products you have published.
    curl -sL "https://www.example.com/index.php?option=com_solidshop&view=sitemap&format=xml" | grep -c "<sitemap>"
    curl -sL "https://www.example.com/index.php?option=com_solidshop&view=sitemap&format=xml&type=products&from=1&to=1000" | grep -c "<loc>"
  3. Compare one entry with its page. Take a product URL from the file and print the canonical the page declares. The two must be identical.
    curl -s "https://www.example.com/shop/oak-desk" | grep -i -o '<link[^>]*rel="canonical"[^>]*>'
  4. Look for a page that should be missing. Search the files for your cart or for an article you set to noindex.
  5. Return to the Sitemaps report after a few days and read the number of URLs discovered.

On a Solidshop store, System → Overview shows the exact sitemap URL and what it includes, and warns about three things: a robots.txt that does not reference it, a sitemap.xml file in the way, and a Global Configuration that still tells search engines not to index the site.

The third file

robots.txt says where a crawler may go, the sitemap says what is there, and Making a Joomla Store Visible to AI Assistants covers what it reads on arrival. Solidshop serves the sitemap from the free core, so the only line you write is the one in robots.txt. Get the core from the extensions page, or open the go-live checklist if your store is already running.