Joomla robots.txt for an Online Store

What the stock file blocks, the old rules that still break migrated sites, how crawlers really read it, the lines a shop should add, and a complete file to copy at the end.

robots.txt is a plain-text file of requests to well-behaved crawlers about which paths to fetch and which to leave alone. It has to sit at the root of the domain, at /robots.txt, in lower case. Joomla’s own file opens with the warning: on a site installed in a subfolder, move the file to the domain root and prefix every path with the folder name.

Two things it is not. It is not access control: the file is public, and RFC 9309, the standard that defines it, says plainly that the rules are not a form of access authorisation. And it is not a way to keep a page out of search. Google’s own introduction says a disallowed page can still be indexed if other sites link to it. Worse, a crawler that is not allowed to fetch a page never sees the noindex tag on it, a point that decides two of the questions below.

What Joomla ships

This is the rule list in Joomla’s robots.txt.dist, identical in the 5.4 and 6.1 branches at the time of writing:

User-agent: *
Disallow: /administrator/
Disallow: /api/
Disallow: /bin/
Disallow: /cache/
Disallow: /cli/
Disallow: /components/
Disallow: /includes/
Disallow: /installation/
Disallow: /language/
Disallow: /layouts/
Disallow: /libraries/
Disallow: /logs/
Disallow: /modules/
Disallow: /plugins/
Disallow: /tmp/

Every line is a system folder no shopper should land on, so nothing in the list needs removing. Note what is absent: /media/, /templates/ and /images/ are all open, because crawlers need the CSS, JavaScript and pictures to see the page as a visitor does.

The .dist suffix matters. The installer renames robots.txt.dist to robots.txt only when no robots.txt exists yet. From then on the file is yours: Joomla updates refresh the .dist copy and leave your edited file alone. A new stock rule therefore never reaches an existing site on its own, so compare the two files after a major upgrade.

The rule that used to break Joomla sites

That last point has a history. Older Joomla releases disallowed /media/ and /templates/, and the 2.5 series blocked /images/ as well. With the stylesheets and scripts off limits, a search engine renders an unstyled, half-working page and judges the site on that. Joomla 3.3 removed the lines, but only in robots.txt.dist; Joomla still carries a post-update message asking site owners to make the change by hand.

If your site began life on Joomla 2.5 or early 3.x and has been migrated since, open the file and look. For a store, Disallow: /images/ is the costly one: it is where product photos usually live. Delete all three lines if they are there.

How crawlers read the file

Three rules from RFC 9309 account for most of the surprises.

A crawler obeys one group, not all of them. It looks for a User-agent group that names it, and falls back to the * group only when none does. So a group written for one bot replaces the general rules for that bot. Add User-agent: Googlebot with a single Disallow beneath it, and Googlebot stops reading your fifteen stock lines.

The longest matching path wins. When an Allow and a Disallow both match a URL, the rule with the longer path is used, whatever order the lines come in. On a tie, Allow wins. This is what makes it possible to reopen one route inside a blocked folder.

Paths are case-sensitive prefixes. Matching starts at the first character of the path, and Disallow: /Tmp/ does not cover /tmp/. Two special characters are defined: * stands for any run of characters, and $ anchors the end of the URL. The User-agent name itself is matched without regard to case.

What a store adds

The Sitemap line

Sitemap: https://www.example.com/index.php?option=com_solidshop&view=sitemap&format=xml

The one addition every site should make. It must be a full URL including the scheme and host, it belongs to no User-agent group so it can go anywhere in the file, and it may appear more than once. Submitting a sitemap in Google Search Console tells Google; this line tells every crawler that reads the file, no account needed. The URL above is the one a Solidshop store serves. With another cart or a sitemap extension, use the address it gives you.

One line is enough only if that sitemap covers the whole site. A shop is more than its catalogue: the About page, the delivery and returns articles and the blog are Joomla content, and a sitemap of products alone leaves them out. From version 1.6.0 the Solidshop sitemap covers both. The Solidshop - Site Content Sitemap plugin, part of the free core and enabled on install, adds the site’s menu items, articles and article categories to the same sitemap index, leaving out anything set to noindex, so the single line above serves the whole site with no second sitemap extension to run. The go-live checklist lists exactly what is included. If your catalogue and your articles come from two different generators, give each its own Sitemap: line.

The API route

Disallow: /api/
Allow: /api/index.php/v1/solidshop/

Stock Disallow: /api/ is right for most sites, since Joomla’s web services are normally for authenticated clients. A store that publishes a read-only catalogue under /api/ and points to it from its llms.txt is inviting crawlers to a door the file tells them is shut. Keep the Disallow and add an Allow for the shop routes only. It is the longer path, so it wins. The rest of that story is in Making a Joomla Store Visible to AI Assistants.

Cart, checkout and account pages

The instinct is to disallow them. For most stores the right answer is to do nothing. These pages should carry a noindex robots meta tag, and a crawler has to be allowed to fetch a page to see it. Disallow the path and the bare URL can still turn up in results through a stray link. Solidshop marks its cart, checkout, account and wishlist pages noindex, follow, so nothing about them belongs in robots.txt. On another cart, view the page source and check before deciding.

Filtered and sorted listing URLs

A category page that can be sorted and filtered is one page with a great many URLs, and on a large catalogue crawlers spend their visits on the variants. The first answer is on the page, not in this file: a canonical link tells the crawler which URL is the real one, and a noindex on deep filter combinations keeps them out of the index. Solidshop does both by default; a listing with two or more filters applied is served noindex, follow.

The blunt answer is a wildcard rule such as Disallow: /*?*sort=, with the parameter name your cart actually uses. It saves crawl effort, at the price described at the top: a blocked URL cannot show its canonical or its noindex. Reach for it only when server logs show crawlers lost in the variants.

AI crawlers: a decision, not a default

The AI companies send up to three bots each, for three different jobs. These are the names their own documentation lists today:

Training crawlers
GPTBot (OpenAI) and ClaudeBot (Anthropic) collect content that may be used to train future models. Google-Extended belongs here too, though it is a token rather than a crawler: Google fetches with its usual user agents, and the token controls whether what it fetched may be used for Gemini training and grounding.
Search-index crawlers
OAI-SearchBot (OpenAI), Claude-SearchBot (Anthropic) and PerplexityBot (Perplexity) build the indexes the assistants search when a question needs the web. Perplexity states that PerplexityBot is not used for training.
User-triggered fetchers
ChatGPT-User, Claude-User and Perplexity-User fetch a page because a person asked about it just now. OpenAI says robots.txt rules may not apply to these requests, and Perplexity says its fetcher generally ignores them, on the reasoning that a person, not a crawler, started the visit.

The sources are OpenAI’s crawler page, Anthropic’s crawler article, Perplexity’s crawler guide and Google’s list of common crawlers. The names change, so check those pages rather than any list, this one included.

Blocking training and blocking search are separate decisions. A store that wants to be named in an assistant’s answer must leave the search-index crawlers alone. A file written when training was the only worry often blocks every AI user agent its author could name, and shuts the shop out of both. On Google the split is cleaner still. Google-Extended has no effect on inclusion or ranking in Google Search, and the AI features inside Search follow the ordinary Googlebot rules.

Whether to allow training is your call. Going by the vendors’ own descriptions, refusing it does not remove you from their search indexes.

A file to copy

The stock rules, the two additions, and an optional group, commented out, that refuses one training crawler.

# robots.txt for a Joomla store. Must live at the domain root.

User-agent: *
Disallow: /administrator/
Disallow: /api/
# Reopen only the public catalogue routes. Longer path than the line above, so it wins.
Allow: /api/index.php/v1/solidshop/
Disallow: /bin/
Disallow: /cache/
Disallow: /cli/
Disallow: /components/
Disallow: /includes/
Disallow: /installation/
Disallow: /language/
Disallow: /layouts/
Disallow: /libraries/
Disallow: /logs/
Disallow: /modules/
Disallow: /plugins/
Disallow: /tmp/

# Optional: refuse one training crawler. Search crawlers are not affected.
# A named group replaces the * group for that crawler.
# User-agent: GPTBot
# Disallow: /

# Full URL required. Tells every crawler where the list of pages is.
Sitemap: https://www.example.com/index.php?option=com_solidshop&view=sitemap&format=xml

Change the host name, and on a cart other than Solidshop drop the Allow line and use your own sitemap address.

Check your work

  1. Read the file as a crawler gets it. This catches a file that was edited in a subfolder or never uploaded.
    curl -s https://www.example.com/robots.txt
  2. Open the robots.txt report in Google Search Console, under the property’s settings. It shows when Google last fetched the file and how many parsing issues it found, and lets you request a recrawl. Google otherwise refreshes its cached copy about once a day.
  3. Fetch one product page and one catalogue route and confirm both answer 200.
    curl -s -o /dev/null -w "%{http_code}\n" https://www.example.com/api/index.php/v1/solidshop/products

On a Solidshop store, System → Overview reads the live file for you and flags three things: a missing Sitemap: reference, a blocked catalogue API path, and AI user agents that have been shut out. The go-live checklist walks through each fix.

One file of four

robots.txt only decides whether a crawler may come in. What it finds once inside is the rest of the work: llms.txt, structured data and a catalogue endpoint, all covered in Making a Joomla Store Visible to AI Assistants. Solidshop ships those three in the free core, so this file is the only one you edit by hand. Get the core from the extensions page, or go straight to the AI discovery steps of the go-live checklist if your store is already running.