Google-PhysicalWeb: Robots.txt & Crawl Policy Reference
Technical reference for the historical Google-PhysicalWeb User-Agent label, with explicit limits around current service status and crawler verification.
AI Summary:
Mozilla/5.0 (Google-PhysicalWeb)is a historical community-registry User-Agent label. Two direct Google Developer Physical Web URLs returned 404 during review, and no current Google source verified this exact token, active service, IP ranges, rate limits, or robots behavior. Treat matching traffic as unverified until current logs and an authoritative source establish more evidence.
Role and policy boundary
The registry records Google-PhysicalWeb as a Google Physical Web crawler and gives this User-Agent:
Mozilla/5.0 (Google-PhysicalWeb)
That record is not a current Google policy. The linked crawler-user-agents repository is a community-maintained syntactic list, and the two Google Developer URLs checked for Physical Web documentation returned 404. The missing pages do not prove that the product never existed or that no related Google service operates; they do mean that this review found no current first-party confirmation of the label.
Treat Physical Web discovery as a historical purpose hypothesis only. A request for a public beacon-related URL, icon, manifest, or HTML page could have many explanations and does not identify the caller. Do not infer current activity, Google ownership, search indexing, AI-training use, IP ranges, or permission to access private data from the string Google-PhysicalWeb.
If logs confirm the exact historical token and your site policy is to exclude it from discovery, a narrow rule could be:
User-agent: Google-PhysicalWeb
Disallow: /
For selective access to public material:
User-agent: Google-PhysicalWeb
Allow: /public/
Allow: /.well-known/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
These are site-owner examples, not recovered Google instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No current Google source reviewed for this profile published a Physical Web crawler network, reverse-DNS rule, rate limit, or product-specific robots policy. Do not apply Googlebot verification rules to this historical token without a current source that expressly covers it.
Compare the requested path with the Physical Web hypothesis without turning that comparison into attribution. Public HTML, manifests, icons, or /.well-known/ resources may be consistent with many discovery workflows. Private endpoint access, high concurrency, repeated retries, or traffic that ignores your published restrictions may indicate spoofing, abuse, a fork, or a different client. These observations establish impact, not operator identity or downstream use.
Evaluate /robots.txt independently. Confirm that the file is served by the intended host, returns a successful status and text content type, and contains the exact group you choose to publish. Test the full header rather than a broad Google substring. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate a historical Google Physical Web client or secure private paths. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If logs show a repeatable but unverified request token, begin with a report-only rule. Adapt the expression to your WAF provider and label it as a pattern match, not official Google identity:
{
"description": "Review historical Google-PhysicalWeb pattern",
"expression": "lower(http.user_agent) contains \"google-physicalweb\"",
"action": "log"
}
After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:
map $http_user_agent $review_google_physicalweb {
default 0;
~*Google-PhysicalWeb 1;
}
server {
location ~ ^/(admin|account|private|internal|api)/ {
if ($review_google_physicalweb) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not create a Google IP allowlist or reverse-DNS exception from the community label or 404 pages. A User-Agent is easy to spoof, and a broad Google match can break legitimate search or monitoring traffic. Test public pages, manifests, icons, sitemaps, media, /.well-known/ resources, account flows, APIs, and approved integrations separately. Pair edge rules with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for the exact Google-PhysicalWeb token and record source IPs, ASNs, PTR results, forward lookups, paths, methods, response sizes, statuses, timing, and rate. Check whether the request source can be verified through a current Google-owned document before assigning provider identity.
Re-check the Google Developer documentation and the community registry when a current source is available. Keep this profile at legacy-label until Google publishes a current canonical User-Agent, purpose, source-verification method, or robots guidance. Treat the two 404 responses as source limitations, not proof of current inactivity.
Decide whether the objective is to preserve public discovery, limit asset requests, protect private material, or reduce crawl load. Publish exact robots rules and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent logs and response behavior.
References
- Community crawler-user-agents repository — community-maintained registry linked by the bot inventory; it is not a Google-owned policy source.
- Google Physical Web documentation URL — returned Google Developer
404 Page Not Foundduring headless-browser review on 2026-08-25. - Google Physical Web documentation alternative — returned Google Developer
404 Page Not Foundduring the same review. - Google crawler overview — general Google verification principles; they should not be applied to this historical label without an applicable first-party source.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.