← Bot Directory/Google-Read-Aloud
Bot directory / search-engine

Google-Read-Aloud: Robots.txt & Crawl Policy Reference

Technical reference for the Google-Read-Aloud User-Agent label, with Google crawler verification guidance and explicit limits around product-specific policy.

AI Summary: Google-Read-Aloud is a community-registry User-Agent label associated with a possible Google read-aloud fetch. Google’s current crawler documentation explains the difference between automatic crawlers and user-triggered fetchers and requires verification through User-Agent, source IP, and reverse DNS. It does not, in the reviewed page, publish a dedicated current policy for this exact token, so treat the header as a lead rather than authentication.

Role and policy boundary

The inventory records this User-Agent pattern:

configuration / code
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/41.0.2272.118 Safari/537.36 (compatible; Google-Read-Aloud; +https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers)

The linked URL now redirects to Google’s Crawling infrastructure overview. Google explains that crawlers automatically discover and scan websites, while fetchers typically make a single request on behalf of a user. It separates common crawlers, special-case crawlers, and user-triggered fetchers. That framework is relevant to a read-aloud label, but the reviewed page did not establish a dedicated current Google-Read-Aloud policy or prove that every matching request belongs to a current Google product.

Treat a public-page read request as a product-specific or user-triggered hypothesis. A matching header does not grant access to private content, bypass authentication, defeat a paywall, or override application authorization. Do not infer AI-training use, general Google Search indexing, current activity, or downstream use from the name.

If your logs confirm this exact token and your business decision is to prevent the read-aloud fetch from discovering selected public paths, publish a narrow group:

configuration / code
User-agent: Google-Read-Aloud
Disallow: /private/
Disallow: /members/
Disallow: /preview/

For complete exclusion:

configuration / code
User-agent: Google-Read-Aloud
Disallow: /

These are site-owner choices, not recovered product instructions. Robots.txt is advisory and cannot protect private or licensed material; use authentication, authorization, signed URLs, and origin controls for those boundaries. Blocking this label does not necessarily block Googlebot or other Google clients.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS hostname, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Google’s current overview says its crawlers and fetchers identify themselves through the HTTP User-Agent, source IP address, and reverse DNS hostname. Use all three signals together.

For a claimed Google request, follow Google’s current verification instructions: perform a reverse DNS lookup on the source IP, assess the hostname against Google’s documented naming patterns, then perform a forward lookup to confirm that the hostname resolves back to the original address. The User-Agent alone is not authentication, and a community or historical token must not inherit Googlebot trust automatically.

Classify the traffic by observable behavior. A one-off request for public article text may be consistent with a user-triggered reading function; a broad crawl, private endpoint access, high concurrency, repeated retries, or unexpected downloads requires investigation. These observations establish impact and operational risk, not identity or downstream use.

Evaluate /robots.txt independently. Confirm that the response is served by the correct host, has a successful status and text content type, and contains the exact Google-Read-Aloud group you intend to publish. Test the full header rather than a broad Google match. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives do not authenticate a read-aloud fetcher or secure private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

Use a report-only rule while confirming whether the traffic is actually associated with a Google product. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe Google-Read-Aloud candidate requests",
  "expression": "lower(http.user_agent) contains \"google-read-aloud\"",
  "action": "log"
}

After independent verification and a deliberate access decision, scope enforcement to selected routes:

configuration / code
map $http_user_agent $block_google_read_aloud {
    default 0;
    ~*Google-Read-Aloud 1;
}

server {
    location ~ ^/(private|members|preview|admin)/ {
        if ($block_google_read_aloud) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and may interfere with an authorized accessibility or preview workflow. Do not create an IP allowlist from a stale third-party list or broadly block all Google clients. Test article pages, canonical metadata, structured data, robots.txt, sitemaps, media, preview routes, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search logs for the complete registered header and any simpler Google-Read-Aloud form. Record source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, and redirect chain. Confirm whether the request is a single user-triggered fetch or an automatic crawl before assigning a policy.

Check that any robots group is exact and that private routes are protected by application authorization. Use noindex or X-Robots-Tag for discovery preferences, but do not mistake them for access control. If you restrict the token, verify that Googlebot and other intentionally allowed identities remain unaffected.

Re-check Google’s crawler overview and verification guidance when the product publishes a dedicated identity or changes its User-Agent. Keep this profile at partially-documented until a current first-party page confirms the exact token and product-specific behavior. Do not claim successful blocking from configuration alone; verify subsequent logs and response behavior.

References

  1. Google crawlers and fetchers overview — Google’s current documentation for crawler categories, fetchers, technical properties, and three-part verification; reviewed with the headless browser on 2026-08-25.
  2. Google crawler verification guidance — use current Google instructions for reverse and forward DNS checks.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.