← Bot Directory/Google-Safety
Bot directory / search-engine

Google-Safety: Robots.txt & Crawl Policy Reference

Technical reference for the community-recorded Google-Safety User-Agent label, with explicit limits around purpose, current activity, and Google verification.

AI Summary: Google-Safety is a community-recorded User-Agent label, not an identity verified by the current Google crawler overview reviewed for this profile. The registry points to a general Google documentation page, but that page does not contain the exact token or publish a dedicated purpose, IP range, rate policy, or robots rule. Treat matching traffic as unverified until logs and a current first-party source establish more evidence.

Role and policy boundary

The registry records the following complete header:

configuration / code
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.5735.179 Mobile Safari/537.36 (compatible; Google-Safety; +http://www.google.com/bot.html)

It describes the label as a Google safety web crawler, but the linked Google Crawling infrastructure overview does not contain the exact Google-Safety token. Google’s documentation does explain that its clients can be common crawlers, special-case crawlers, or user-triggered fetchers. That general framework cannot establish which category this unverified label belongs to.

Do not infer a safety-scanning purpose, Google ownership, current activity, search indexing, AI-training use, or permission to access private content from the name. A request to a public page, a security-related path, or a policy document may have many explanations. A User-Agent is a claimed string supplied by the requester, not authentication.

If your logs confirm this exact token and your site policy is to exclude it from discovery, a narrow rule could be:

configuration / code
User-agent: Google-Safety
Disallow: /

For selective access:

configuration / code
User-agent: Google-Safety
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

These are site-owner examples, not published Google instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries. Do not assume this group will affect Googlebot or another verified Google product.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Google’s general documentation says its crawlers and fetchers identify themselves through the HTTP User-Agent, source IP address, and reverse DNS hostname. Apply those signals to the actual request, not to the registry label.

For a claimed Google request, follow Google’s current verification instructions: reverse-resolve the source IP, assess the hostname using Google’s published guidance, then forward-resolve the hostname to confirm that it maps back to the original address. Do not treat an arbitrary PTR, a google.com substring, or the registered header alone as proof of identity. If the source cannot be independently verified, keep it classified as an unknown client.

Classify observed behavior without turning it into an attribution. Low-volume access to public pages may be consistent with many inspection or monitoring tools. High-concurrency traversal, private endpoint access, repeated retries, unusual downloads, or traffic that ignores your published restrictions may indicate spoofing, abuse, a fork, or a different client. These observations establish operational impact, not operator identity or downstream use.

Evaluate /robots.txt independently. Confirm that the response is served by the intended host, returns a successful status and text content type, and contains the exact Google-Safety group you intentionally publish. Test the full token and any global group separately. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives do not authenticate a safety service or secure private paths. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

When the identity and purpose are unverified, begin with a report-only rule. Adapt this expression to your WAF provider and label it as a candidate pattern rather than official Google traffic:

configuration / code
{
  "description": "Review Google-Safety candidate requests",
  "expression": "lower(http.user_agent) contains \"google-safety\"",
  "action": "log"
}

After independent verification and a deliberate access decision, scope enforcement to sensitive routes:

configuration / code
map $http_user_agent $review_google_safety {
    default 0;
    ~*Google-Safety 1;
}

server {
    location ~ ^/(admin|account|private|internal|api)/ {
        if ($review_google_safety) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Do not block all Google traffic, invent an IP allowlist, or reuse Googlebot exceptions. A User-Agent is easy to spoof and a broad expression can break legitimate search, Ads, inspection, or monitoring clients. Test public HTML, security documents, sitemaps, media, account flows, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search logs for the exact Google-Safety token and preserve representative source IPs, ASNs, PTR results, forward lookups, paths, methods, response sizes, statuses, timing, and rate. Check whether a current Google-owned document identifies the client before assigning Google identity or a safety-scanning role.

Verify that any robots group is exact and that private routes are protected by application authorization. Use noindex or X-Robots-Tag for discovery preferences, but do not mistake them for access control. If you restrict the token, confirm that Googlebot and other intentionally allowed identities remain unaffected.

Re-check Google’s crawler overview and specific-crawler pages when a current record becomes available. Keep this profile at partially-documented until Google publishes a current canonical User-Agent, purpose, source-verification method, or robots guidance. Do not claim successful blocking from configuration alone; verify subsequent logs and response behavior.

References

  1. Google crawlers and fetchers overview — Google’s current general documentation for crawler categories and three-part verification; the exact Google-Safety token was not present during review on 2026-08-25.
  2. Google crawler verification guidance — use current Google instructions for reverse and forward DNS checks.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.