← Bot Directory/Storebot-Google
Bot directory / search-engine

Storebot-Google: Robots.txt & Crawl Policy Reference

Technical reference for Google StoreBot, the e-commerce crawler used for Google Shopping, including its User-Agent patterns, robots.txt compliance, and DNS verification.

AI Summary: Storebot-Google (Google StoreBot) is an official Google crawler used to discover and index product pages for all surfaces of Google Shopping. Google’s documentation classifies it as a common crawler that strictly obeys robots.txt rules using the Storebot-Google token. It publishes both desktop and mobile User-Agent patterns. Verify requests using Google’s standard reverse/forward DNS method or its common-crawlers.json IP list before trusting the header.

Role and policy boundary

Google StoreBot is deployed to index product and e-commerce pages. Crawling preferences addressed to this crawler affect all surfaces of Google Shopping, including the Shopping tab in Google Search.

Google publishes two User-Agent patterns for this crawler. The desktop agent is:

configuration / code
Mozilla/5.0 (X11; Linux x86_64; Storebot-Google/1.0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Safari/537.36

The mobile agent is:

configuration / code
Mozilla/5.0 (Linux; Android 8.0; Pixel 2 Build/OPD3.170816.012; Storebot-Google/1.0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36

In these patterns, W.X.Y.Z represents the Chromium release version currently used by the crawler. The registry records a static Chrome version (79.0.3945.88), but expect the observed version to update over time.

As a Google common crawler, Storebot-Google always respects robots.txt rules for automatic crawls. For selective access to public product pages:

configuration / code
User-agent: Storebot-Google
Allow: /products/
Allow: /catalog/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

For a complete exclusion from Google Shopping indexing:

configuration / code
User-agent: Storebot-Google
Disallow: /

These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. The product-indexing purpose does not establish permission for third-party model training or generative AI input outside of Google's specified shopping surfaces.

Layered verification

Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs.

Because a User-Agent is easily spoofed, do not treat the header alone as authentication. Use Google’s standard network verification: perform a reverse DNS lookup on the source IP. For common crawlers, the hostname will match crawl-***-***-***-***.googlebot.com or geo-crawl-***-***-***-***.geo.googlebot.com. Then perform a forward lookup to confirm it resolves back to the same IP. Alternatively, verify the IP against Google’s published common-crawlers.json list. A Storebot-Google header from an unverified IP or a residential network is a spoof and should be blocked.

Compare observed traffic with the documented product-indexing role. Public HTML, product metadata, feeds, sitemaps, and ordinary assets are consistent with discovery. Private endpoints, account routes, APIs, licensed content, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not permission.

Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact Storebot-Google group you intend to publish. Check precedence, path matching, and actual request behavior. Meta directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

Begin in report-only mode and correlate the Storebot-Google User-Agent with reverse/forward DNS verification or the Google IP JSON. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe Storebot-Google candidates before enforcement",
  "expression": "lower(http.user_agent) contains \"storebot-google\"",
  "action": "log"
}

For a deliberately restricted private route, use the exact token only after verifying the source network:

configuration / code
map $http_user_agent $block_storebot_private {
    default 0;
    ~*Storebot-Google 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_storebot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

This rule is route-scoped and does not prove identity. Fetch Google’s JSON list dynamically if you implement a network-level allowlist. Do not invent additional ASN exceptions, rates, or training permissions without operator documentation.

A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact Storebot-Google string, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Perform reverse and forward DNS verification to confirm a googlebot.com hostname, or check the IP against common-crawlers.json, before trusting the client. Treat requests from unverified networks as spoofed.

Check your robots file for an exact Storebot-Google group. Test precedence and path behavior, and verify actual logs after publishing. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.

Review product indexing, AI input, reference use, and model-training decisions separately. Google explicitly states the crawler is used for Google Shopping surfaces.

Keep the profile at documented: the operator publishes the exact User-Agent patterns, robots identifier, product-indexing purpose, and network verification methods. Re-check the crawler documentation when the source network or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.

References

  1. Google's common crawlers — official documentation detailing Storebot-Google, its User-Agent patterns, and robots.txt token; reviewed with the headless browser on 2026-08-25.
  2. Overview of Google crawlers and fetchers — context on Google's crawling infrastructure and network verification.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.