Google-InspectionTool: Robots.txt & Crawl Policy Reference
Technical reference for Google's InspectionTool crawler, including published User-Agent examples, verification layers, robots considerations, and safe edge controls.
AI Summary: Google documents
Google-InspectionToolas a crawler/fetcher identity used for testing and debugging search indexing. Published examples includeGoogle-InspectionTool/1.0, including mobile and non-mobile forms. Verify requests with the complete User-Agent, source IP, and reverse DNS; do not treat the header alone as authentication, and distinguish an on-demand inspection from an automatic crawl.
Role and policy boundary
Google’s crawler documentation lists Google-InspectionTool among the User-Agent identities used by its crawling infrastructure. The registry’s mobile example is:
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Google-InspectionTool/1.0)
Google also documents a simpler form:
Mozilla/5.0 (compatible; Google-InspectionTool/1.0)
The stated role is testing and debugging search indexing. It is not the same identity as Googlebot, Googlebot-Image, Storebot-Google, or a generic Google fetcher. Keep controls specific to Google-InspectionTool; do not use a broad Google substring that could affect unrelated products.
Google’s documentation distinguishes common crawlers, special-case crawlers, and user-triggered fetchers. Inspection traffic may be initiated by a product tool or a user request, so interpret the requested URL, timing, and product context alongside the header. A matching header does not grant access to private content and does not prove that the request should bypass authentication, paywalls, rate limits, or origin controls.
If you want to prevent this inspection identity from discovering a public path, publish a narrow robots group after considering the debugging and indexing consequences:
User-agent: Google-InspectionTool
Disallow: /private/
Disallow: /internal/
Disallow: /preview/
For a complete exclusion:
User-agent: Google-InspectionTool
Disallow: /
Robots.txt is advisory and cannot secure private or licensed material. Use authentication, authorization, signed URLs, and origin controls for those boundaries. Do not assume that blocking an inspection request blocks Google Search crawling by other identities.
Layered verification
Begin with access logs and retain the complete User-Agent, source IP, ASN, reverse-DNS hostname, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Google’s documentation says its crawlers and fetchers identify themselves through the HTTP User-Agent, source IP address, and reverse DNS hostname. The three signals should be evaluated together.
Verify a claimed Google source by performing a reverse DNS lookup on the requesting IP, checking that the hostname is within a Google-controlled naming pattern documented by Google, then performing a forward lookup to ensure the hostname resolves back to the original IP. Use Google’s current verification instructions for the exact product and keep the checks current; never treat an arbitrary PTR record or User-Agent as sufficient.
Classify the request by observable role. A low-volume request for a public page during an inspection workflow may be consistent with debugging. Repeated high-rate traversal, private endpoint access, unexpected asset downloads, or traffic that does not match independently verified Google infrastructure requires investigation. These observations establish impact and operational risk, not a right to crawl or a downstream-use claim.
Evaluate /robots.txt independently. Confirm that it is served by the correct host, returns a successful status and text content type, and contains the exact Google-InspectionTool group you intend to publish. Test both mobile and non-mobile User-Agent examples and confirm how your parser handles a global group. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives do not authenticate Google, secure private paths, or guarantee that every product-specific fetcher will behave identically. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
Use a report-only rule first if you are diagnosing unexpected inspection traffic. Adapt this expression to your WAF provider and pair it with IP/reverse-DNS evidence before making an attribution:
{
"description": "Observe Google InspectionTool requests",
"expression": "lower(http.user_agent) contains \"google-inspectiontool\"",
"action": "log"
}
For a sensitive route control, use a narrow token match and avoid blocking all Google agents:
map $http_user_agent $block_google_inspection_tool {
default 0;
~*Google-InspectionTool 1;
}
server {
location ~ ^/(private|internal|preview|admin)/ {
if ($block_google_inspection_tool) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof and can also interfere with an authorized inspection workflow. Do not create an IP allowlist from memory or from stale third-party lists; follow Google’s current verification documentation. Test public HTML, canonical metadata, structured data, robots.txt, sitemaps, media, preview routes, accounts, APIs, and approved integrations separately. Pair edge rules with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for both the complete mobile form and the non-mobile Google-InspectionTool/1.0 form. Record source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, and redirect chain. Confirm whether the request aligns with an on-demand inspection or another documented Google product before labeling it as an official request.
Check that your robots group is exact and that any restricted route is also protected by application authorization. Use noindex or X-Robots-Tag for discovery preferences, but do not mistake them for access control. If you block inspection traffic, verify subsequent logs and confirm that Googlebot and other intentionally allowed identities remain unaffected.
Re-check Google’s crawler overview and verification instructions when Google changes its product identities or network guidance. Do not claim successful blocking from configuration alone; verify response behavior and search-debugging impact with an authorized test.
References
- Google crawlers and fetchers overview — Google’s first-party documentation for crawler categories, User-Agent examples, technical properties, and verification signals; reviewed from a headless-browser capture dated 2026-08-25.
- Google crawler verification guidance — use current Google instructions for reverse and forward DNS verification rather than relying on User-Agent alone.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.