What the stock file blocks, the old rules that still break migrated sites, how crawlers really read it, the lines a shop should add, and a complete file to copy at the end.
robots.txt is a plain-text file of requests to
well-behaved crawlers about which paths to fetch and which to leave
alone. It has to sit at the root of the domain, at
/robots.txt, in lower case. Joomla’s own file
opens with the warning: on a site installed in a subfolder, move
the file to the domain root and prefix every path with the folder
name.
Two things it is not. It is not access control: the file is public,
and
RFC
9309, the standard that defines it, says plainly that the rules
are not a form of access authorisation. And it is not a way to keep
a page out of search. Google’s
own
introduction says a disallowed page can still be indexed if
other sites link to it. Worse, a crawler that is not allowed to
fetch a page never sees the noindex tag on it, a point
that decides two of the questions below.
What Joomla ships
This is the rule list in Joomla’s robots.txt.dist,
identical in the 5.4 and 6.1 branches at the time of writing:
User-agent: *
Disallow: /administrator/
Disallow: /api/
Disallow: /bin/
Disallow: /cache/
Disallow: /cli/
Disallow: /components/
Disallow: /includes/
Disallow: /installation/
Disallow: /language/
Disallow: /layouts/
Disallow: /libraries/
Disallow: /logs/
Disallow: /modules/
Disallow: /plugins/
Disallow: /tmp/
Every line is a system folder no shopper should land on, so nothing
in the list needs removing. Note what is absent: /media/,
/templates/ and /images/ are all open,
because crawlers need the CSS, JavaScript and pictures to see the
page as a visitor does.
The .dist suffix matters. The installer renames
robots.txt.dist to robots.txt only when
no robots.txt exists yet. From then on the file is
yours: Joomla updates refresh the .dist copy and leave
your edited file alone. A new stock rule therefore never reaches an
existing site on its own, so compare the two files after a major
upgrade.
The rule that used to break Joomla sites
That last point has a history. Older Joomla releases disallowed
/media/ and /templates/, and the 2.5
series blocked /images/ as well. With the stylesheets
and scripts off limits, a search engine renders an unstyled,
half-working page and judges the site on that. Joomla 3.3 removed
the lines, but only in robots.txt.dist; Joomla still
carries a post-update message asking site owners to make the
change by hand.
If your site began life on Joomla 2.5 or early 3.x and has been
migrated since, open the file and look. For a store,
Disallow: /images/ is the costly one: it is where
product photos usually live. Delete all three lines if they are
there.
How crawlers read the file
Three rules from RFC 9309 account for most of the surprises.
A crawler obeys one group, not all of them. It
looks for a User-agent group that names it, and falls
back to the * group only when none does. So a group
written for one bot replaces the general rules for that
bot. Add User-agent: Googlebot with a single
Disallow beneath it, and Googlebot stops reading your
fifteen stock lines.
The longest matching path wins. When an
Allow and a Disallow both match a URL,
the rule with the longer path is used, whatever order the lines
come in. On a tie, Allow wins. This is what makes it
possible to reopen one route inside a blocked folder.
Paths are case-sensitive prefixes. Matching
starts at the first character of the path, and
Disallow: /Tmp/ does not cover /tmp/.
Two special characters are defined: * stands for any
run of characters, and $ anchors the end of the URL.
The User-agent name itself is matched without regard
to case.
What a store adds
The Sitemap line
Sitemap: https://www.example.com/index.php?option=com_solidshop&view=sitemap&format=xml
The one addition every site should make. It must be a full URL
including the scheme and host, it belongs to no
User-agent group so it can go anywhere in the file,
and it may appear more than once. Submitting a sitemap in Google
Search Console tells Google; this line tells every crawler that
reads the file, no account needed. The URL above is the one a
Solidshop store serves. With another cart or a sitemap extension,
use the address it gives you.
One line is enough only if that sitemap covers the whole site. A
shop is more than its catalogue: the About page, the delivery and
returns articles and the blog are Joomla content, and a sitemap of
products alone leaves them out. From version 1.6.0 the Solidshop
sitemap covers both. The
Solidshop - Site Content Sitemap plugin, part of
the free core and enabled on install, adds the site’s menu
items, articles and article categories to the same sitemap index,
leaving out anything set to noindex, so the single
line above serves the whole site with no second sitemap extension
to run. The go-live checklist
lists exactly what is included. If your catalogue and your
articles come from two different generators, give each its own
Sitemap: line.
The API route
Disallow: /api/
Allow: /api/index.php/v1/solidshop/
Stock Disallow: /api/ is right for most sites, since
Joomla’s web services are normally for authenticated
clients. A store that publishes a read-only catalogue under
/api/ and points to it from its
llms.txt is inviting crawlers to a door the file
tells them is shut. Keep the
Disallow and add an Allow for the shop
routes only. It is the longer path, so it wins. The rest of that
story is in
Making a Joomla Store Visible to AI
Assistants.
Cart, checkout and account pages
The instinct is to disallow them. For most stores the right
answer is to do nothing. These pages should carry a
noindex robots meta tag, and a crawler has to be
allowed to fetch a page to see it. Disallow the path and the bare
URL can still turn up in results through a stray link. Solidshop
marks its cart,
checkout, account and wishlist pages
noindex, follow, so nothing about them belongs in
robots.txt. On another cart, view the page source and
check before deciding.
Filtered and sorted listing URLs
A category page that can be sorted and filtered is one page with a
great many URLs, and on a large catalogue crawlers spend their
visits on the variants. The first
answer is on the page, not in this file: a canonical link tells
the crawler which URL is the real one, and a
noindex on deep filter combinations keeps them out of
the index. Solidshop does both by default; a listing with two or
more filters applied is served noindex, follow.
The blunt answer is a wildcard rule such as
Disallow: /*?*sort=, with the parameter name your
cart actually uses. It saves crawl effort, at the price described
at the top: a blocked URL cannot show its canonical or its
noindex. Reach for it only when server logs show
crawlers lost in the variants.
AI crawlers: a decision, not a default
The AI companies send up to three bots each, for three different jobs. These are the names their own documentation lists today:
- Training crawlers
-
GPTBot(OpenAI) andClaudeBot(Anthropic) collect content that may be used to train future models.Google-Extendedbelongs here too, though it is a token rather than a crawler: Google fetches with its usual user agents, and the token controls whether what it fetched may be used for Gemini training and grounding. - Search-index crawlers
-
OAI-SearchBot(OpenAI),Claude-SearchBot(Anthropic) andPerplexityBot(Perplexity) build the indexes the assistants search when a question needs the web. Perplexity states thatPerplexityBotis not used for training. - User-triggered fetchers
-
ChatGPT-User,Claude-UserandPerplexity-Userfetch a page because a person asked about it just now. OpenAI saysrobots.txtrules may not apply to these requests, and Perplexity says its fetcher generally ignores them, on the reasoning that a person, not a crawler, started the visit.
The sources are OpenAI’s crawler page, Anthropic’s crawler article, Perplexity’s crawler guide and Google’s list of common crawlers. The names change, so check those pages rather than any list, this one included.
Blocking training and blocking search are separate decisions. A
store that wants to be named in an
assistant’s answer must leave the search-index crawlers
alone. A file written when training was the only worry often
blocks every AI user agent its author could name, and shuts the
shop out of both. On Google the split is cleaner still.
Google-Extended has no effect on inclusion or ranking
in Google Search, and the AI features inside Search follow the
ordinary Googlebot rules.
Whether to allow training is your call. Going by the vendors’ own descriptions, refusing it does not remove you from their search indexes.
A file to copy
The stock rules, the two additions, and an optional group, commented out, that refuses one training crawler.
# robots.txt for a Joomla store. Must live at the domain root.
User-agent: *
Disallow: /administrator/
Disallow: /api/
# Reopen only the public catalogue routes. Longer path than the line above, so it wins.
Allow: /api/index.php/v1/solidshop/
Disallow: /bin/
Disallow: /cache/
Disallow: /cli/
Disallow: /components/
Disallow: /includes/
Disallow: /installation/
Disallow: /language/
Disallow: /layouts/
Disallow: /libraries/
Disallow: /logs/
Disallow: /modules/
Disallow: /plugins/
Disallow: /tmp/
# Optional: refuse one training crawler. Search crawlers are not affected.
# A named group replaces the * group for that crawler.
# User-agent: GPTBot
# Disallow: /
# Full URL required. Tells every crawler where the list of pages is.
Sitemap: https://www.example.com/index.php?option=com_solidshop&view=sitemap&format=xml
Change the host name, and on a cart other than Solidshop drop the
Allow line and use your own sitemap address.
Check your work
-
Read the file as a crawler gets it. This catches a
file that was edited in a subfolder or never uploaded.
curl -s https://www.example.com/robots.txt - Open the robots.txt report in Google Search Console, under the property’s settings. It shows when Google last fetched the file and how many parsing issues it found, and lets you request a recrawl. Google otherwise refreshes its cached copy about once a day.
-
Fetch one product page and one catalogue route
and confirm both answer
200.curl -s -o /dev/null -w "%{http_code}\n" https://www.example.com/api/index.php/v1/solidshop/products
On a Solidshop store, System → Overview reads
the live file for you and flags three things: a missing
Sitemap: reference, a blocked catalogue API path, and
AI user agents that have been shut out. The
go-live checklist walks through
each fix.
One file of four
robots.txt only decides whether a crawler may come
in. What it finds once inside is the rest of the work:
llms.txt, structured data and a catalogue endpoint,
all covered in
Making a Joomla Store Visible to AI
Assistants. Solidshop ships those three in the free core, so
this file is the only one you edit by hand. Get the core from the
extensions page, or go
straight to the AI
discovery steps of the go-live checklist if your store is
already running.