Skip to main content

Robots.txt and Google Shopping: what to block, what to allow

Google’s Merchant Center help asks you to let its user agents reach your product pages and images, so robots.txt must not block them. Block only cart, checkout, account and internal search, then test a real product URL after every change. One copied line from a staging site can close the whole shop.

Critical5 min read

Why this matters

Google's “How to fix: Unable to check product pages” page is written about exactly this file. It asks you to update robots.txt so Google can fetch the landing pages you submit and, to open the whole site, to allow the two Google user agents it names: one for landing pages, one for images. It also tells you where the file lives: “The robots.txt file can usually be found in the root directory of the web server”. Its sister page, “How to fix: Landing page not working”, adds AdsBot-Google to the list and asks you to make sure it “isn't blocked by your robots.txt file”. And “How to fix: Product page unavailable” treats an unreachable file as an error of its own: “The product page was unavailable for Google to check because your online store's robots.txt file couldn't be reached.”

The rules are stricter than they look. Google's robots.txt specification matches from the start of the path: a rule for /fish “Matches any path that starts with /fish. Note that the matching is case-sensitive.” So a hopeful Disallow: /p also shuts /products/ and /pages/. “The field name (disallow) is case-insensitive, but its value is case-sensitive.” And when rules clash, “In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.”

Robots.txt is not a privacy tool either. Google's introduction says it “is not a mechanism for keeping a web page out of Google”, and “A page that's disallowed in robots.txt can still be indexed if linked to from other sites.” Use a noindex tag or a password for pages you want kept out of search. On hosted platforms you may not touch the file at all: “If you use a CMS, such as Wix or Blogger, you might not need to (or be able to) edit your robots.txt file directly.”

The free scan reads your robots.txt and reports a rule that disallows the whole site, for every user agent or for Google's, as a failure. Narrower rules that catch /products/ or /collections/ are yours to test with URL Inspection; the misrepresentation checker lists the access checks the scan runs. For firewalls and bot challenges, which block pages even when robots.txt is clean, see bot protection and Google Shopping.

Typical evidence

One greedy prefix can take the catalogue dark

robots.txt matches by prefix and is blunt. Disallow: /collections/ doesn't hide one thing — it blocks every collection page on the site. Block the plumbing (admin, cart, checkout); leave public product and policy pages open to shoppers and permitted crawlers.

robots.txt — blocks catalogue content
User-agent: *
Disallow: /collections/
Disallow: /admin/
Disallow: /cart
robots.txt — plumbing only, content open
User-agent: *
Disallow: /admin/
Disallow: /cart
Disallow: /checkout
Allow: /products/
Allow: /collections/
Sitemap: https://yourdomain.com/sitemap.xml
Test a real product URL and a policy URL in Search Console's robots report — confirm Google says 'allowed'. The matching rules are fiddlier than they look, so don't eyeball it.

The public signals this check looks for:

  1. A Disallow…

    A Disallow: / line for all user agents, written for a staging site, went live with the rest of the build.

  2. A rule such as Disallow…

    A rule such as Disallow: /collections/ or Disallow: /products/ blocks the exact pages your product links point to.

  3. A short prefix rule catches more than intended…

    A short prefix rule catches more than intended: Disallow: /p meant for one folder also blocks /products/ and /pages/.

  4. A separate group written for one Google user agent carries a stricter rule than the group …

    A separate group written for one Google user agent carries a stricter rule than the group for everyone else, and that group is the one that user agent follows.

  5. Product images sit under a disallowed folder, so the photos cannot be fetched even though …

    Product images sit under a disallowed folder, so the photos cannot be fetched even though the pages can.

  6. The robots

    The robots.txt address itself returns a server error or times out.

What it looks like once it is right

robots.txt blocks /cart, /checkout, /account and /search, lists the sitemap, and nothing else. URL Inspection on a product page, a policy page and a product image shows each one allowed.

Common mistakes

Common mistake

robots.txt reads User-agent: * and Disallow: /p, meant to hide a /private folder. It also blocks every URL under /products/ and /pages/, so the product links and the returns policy are off limits to Google.

Fix checklist

Test the real URLs against the real rules

Use Search Console’s robots report to test how the published rules apply to a specific URL. Test a product URL and a policy URL instead of relying on what the rule appears to say.

robots.txt Tester · Googlebot
/products/oak-desk-lampno matching Disallowallowed
/policies/returnsno matching Disallowallowed
/collections/lightingDisallow: /collections/blocked
The collection URL is blocked by a prefix rule that also kills the products beneath it. Keep the block list to genuine plumbing — admin, cart, checkout — and re-test.

Copy this, and change the parts in your own words

User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /account
Disallow: /search

Sitemap: https://www.example.com/sitemap.xml

Questions merchants ask

Which user agents does robots.txt need to allow for Google Shopping?

Google's “Unable to check product pages” help page names two Google user agents to allow across your whole site, one for landing pages and one for images, and its “Landing page not working” page adds AdsBot-Google. The safe setup is simple: no rule that blocks any of them, and blocks only on cart, checkout, account and search.

Does blocking a page in robots.txt keep it out of Google?

No. Google's introduction to robots.txt says it is not a mechanism for keeping a web page out of Google, and a disallowed page can still be indexed if other sites link to it. Use a noindex tag or a password for pages you want kept out of search.

What happens if my robots.txt file returns an error?

Google's “Product page unavailable” help page lists an unreachable robots.txt as its own error: product pages could not be checked because the file couldn't be reached. Make sure the file loads with a 200 response, or returns a plain 404 if you have no rules at all. Never let it time out.

Remediation

Risk signal

Robots.txt is a tiny file that can switch off a whole catalogue, and it survives migrations because nobody reads it. Read it line by line after every launch, test real URLs, and let a scan catch the site-wide block.
PriorityTreat this and any other highest-severity findings as first-priority work, then document each fix.
EvidenceRecord the current state before each change, apply the fix, then capture the corrected state so every change is evidenced.

Similar cases

Sources

  1. How to fix: Unable to check product pages (robots.txt)Google Merchant Center Help — support.google.com
  2. How Google interprets the robots.txt specificationGoogle Search Central — developers.google.com
  3. Introduction to robots.txtGoogle Search Central — developers.google.com

Last reviewed 23 Sep 2026.

That is one issue. The library documents 134.

The free scan lists what it finds on your store. The paid report adds the affected pages, captured evidence and step-by-step fixes. Start free, with no account needed.