Robots.txt and Google Shopping: what to block, what to allow
Google’s Merchant Center help asks you to let its user agents reach your product pages and images, so robots.txt must not block them. Block only cart, checkout, account and internal search, then test a real product URL after every change. One copied line from a staging site can close the whole shop.
Why this matters
Google's “How to fix: Unable to check product pages” page is written about exactly this file. It asks you to update robots.txt so Google can fetch the landing pages you submit and, to open the whole site, to allow the two Google user agents it names: one for landing pages, one for images. It also tells you where the file lives: “The robots.txt file can usually be found in the root directory of the web server”. Its sister page, “How to fix: Landing page not working”, adds AdsBot-Google to the list and asks you to make sure it “isn't blocked by your robots.txt file”. And “How to fix: Product page unavailable” treats an unreachable file as an error of its own: “The product page was unavailable for Google to check because your online store's robots.txt file couldn't be reached.”
The rules are stricter than they look. Google's robots.txt specification matches from the start of the path: a rule for /fish “Matches any path that starts with /fish. Note that the matching is case-sensitive.” So a hopeful Disallow: /p also shuts /products/ and /pages/. “The field name (disallow) is case-insensitive, but its value is case-sensitive.” And when rules clash, “In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.”
Robots.txt is not a privacy tool either. Google's introduction says it “is not a mechanism for keeping a web page out of Google”, and “A page that's disallowed in robots.txt can still be indexed if linked to from other sites.” Use a noindex tag or a password for pages you want kept out of search. On hosted platforms you may not touch the file at all: “If you use a CMS, such as Wix or Blogger, you might not need to (or be able to) edit your robots.txt file directly.”
The free scan reads your robots.txt and reports a rule that disallows the whole site, for every user agent or for Google's, as a failure. Narrower rules that catch /products/ or /collections/ are yours to test with URL Inspection; the misrepresentation checker lists the access checks the scan runs. For firewalls and bot challenges, which block pages even when robots.txt is clean, see bot protection and Google Shopping.
Typical evidence
One greedy prefix can take the catalogue dark
robots.txt matches by prefix and is blunt. Disallow: /collections/ doesn't hide one thing — it blocks every collection page on the site. Block the plumbing (admin, cart, checkout); leave public product and policy pages open to shoppers and permitted crawlers.
The public signals this check looks for:
A Disallow…
A Disallow: / line for all user agents, written for a staging site, went live with the rest of the build.
A rule such as Disallow…
A rule such as Disallow: /collections/ or Disallow: /products/ blocks the exact pages your product links point to.
A short prefix rule catches more than intended…
A short prefix rule catches more than intended: Disallow: /p meant for one folder also blocks /products/ and /pages/.
A separate group written for one Google user agent carries a stricter rule than the group …
A separate group written for one Google user agent carries a stricter rule than the group for everyone else, and that group is the one that user agent follows.
Product images sit under a disallowed folder, so the photos cannot be fetched even though …
Product images sit under a disallowed folder, so the photos cannot be fetched even though the pages can.
The robots
The robots.txt address itself returns a server error or times out.
What it looks like once it is right
robots.txt blocks /cart, /checkout, /account and /search, lists the sitemap, and nothing else. URL Inspection on a product page, a policy page and a product image shows each one allowed.
Common mistakes
Common mistake
Fix checklist
Test the real URLs against the real rules
Use Search Console’s robots report to test how the published rules apply to a specific URL. Test a product URL and a policy URL instead of relying on what the rule appears to say.
Copy this, and change the parts in your own words
User-agent: * Disallow: /cart Disallow: /checkout Disallow: /account Disallow: /search Sitemap: https://www.example.com/sitemap.xml
Questions merchants ask
Which user agents does robots.txt need to allow for Google Shopping?
Google's “Unable to check product pages” help page names two Google user agents to allow across your whole site, one for landing pages and one for images, and its “Landing page not working” page adds AdsBot-Google. The safe setup is simple: no rule that blocks any of them, and blocks only on cart, checkout, account and search.
Does blocking a page in robots.txt keep it out of Google?
No. Google's introduction to robots.txt says it is not a mechanism for keeping a web page out of Google, and a disallowed page can still be indexed if other sites link to it. Use a noindex tag or a password for pages you want kept out of search.
What happens if my robots.txt file returns an error?
Google's “Product page unavailable” help page lists an unreachable robots.txt as its own error: product pages could not be checked because the file couldn't be reached. Make sure the file loads with a 200 response, or returns a plain 404 if you have no rules at all. Never let it time out.
Remediation
Risk signal
Similar cases
Sources
Last reviewed 23 Sep 2026.
That is one issue. The library documents 134.
The free scan lists what it finds on your store. The paid report adds the affected pages, captured evidence and step-by-step fixes. Start free, with no account needed.