Google-Structured-Data-Testing-Tool: Robots.txt & Crawl Policy Reference
Technical reference for the retired Google Structured Data Testing Tool request label, with current replacement guidance and explicit limits around crawler verification.
AI Summary: Google’s Structured Data Testing Tool is a retired tool label. The registry-linked URL now redirects to Google’s structured-data guidance, which recommends the Rich Results Test and Schema Markup Validator. Treat the recorded User-Agent as a historical request pattern, not proof of current activity; verify any matching traffic before changing robots or WAF rules.
Role and policy boundary
The registry records this historical User-Agent:
Mozilla/5.0 (compatible; Google-Structured-Data-Testing-Tool +https://search.google.com/structured-data/testing-tool)
Google Search Central’s current structured-data page says the old Structured Data Testing Tool was removed and that Google-specific validation should use the Rich Results Test, while generic Schema.org validation should use the Schema Markup Validator. Google’s 2020 update explained the deprecation and migration; later updates stated that the Schema Markup Validator stabilized and the old tool redirected to a landing page.
This establishes a retired tool history, not current crawler operation. A request carrying the historical token could be a legacy client, a stale integration, a test harness, or a spoofed header. Do not infer that it is current Google Search crawling, that it may access private pages, or that it should bypass authentication, paywalls, rate limits, or origin controls.
If logs confirm the exact historical token and you want to exclude it from public discovery, a narrow site-owner rule could be:
User-agent: Google-Structured-Data-Testing-Tool
Disallow: /
For a selective policy while preserving public structured-data test fixtures:
User-agent: Google-Structured-Data-Testing-Tool
Allow: /public/examples/
Allow: /structured-data/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
These are defensive examples, not current Google instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries. Do not assume this historical rule controls the current Rich Results Test or Schema Markup Validator.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. The reviewed structured-data pages did not publish a dedicated current IP range, reverse-DNS procedure, or rate policy for the retired tool label.
For a claimed Google request, use Google’s current general verification guidance: reverse-resolve the source IP, evaluate the hostname against Google’s published naming patterns, then forward-resolve the hostname to confirm that it maps back to the original address. The User-Agent alone is not authentication, and a historical tool string must not inherit Googlebot trust automatically.
Classify the request by observable behavior. A request to a public HTML document or a structured-data example might be consistent with validation, but it does not identify the caller. Private endpoint access, high concurrency, repeated retries, unexpected downloads, or traffic that ignores your published restrictions may indicate a stale client, spoofing, abuse, or another tool. These observations establish impact, not operator identity or downstream use.
Evaluate /robots.txt independently. Confirm that the file is served by the correct host, returns a successful status and text content type, and contains the exact historical group you choose to publish. Test the full token and any global group separately. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate a retired validation tool or secure private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If you see the historical token in logs, begin with a report-only rule and confirm whether it has any legitimate owner or business purpose. Adapt the expression to your WAF provider:
{
"description": "Review retired Structured Data Testing Tool traffic",
"expression": "lower(http.user_agent) contains \"structured-data-testing-tool\"",
"action": "log"
}
After reviewing false positives and confirming the decision, scope enforcement to sensitive routes:
map $http_user_agent $review_retired_structured_data_tool {
default 0;
~*Google-Structured-Data-Testing-Tool 1;
}
server {
location ~ ^/(admin|account|private|internal|api)/ {
if ($review_retired_structured_data_tool) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not broadly block Google clients or create an IP allowlist from the retired label. A User-Agent is easy to spoof and an overbroad expression can disrupt current Rich Results testing, Search Console, authorized monitoring, or legitimate developer tools. Test public HTML, JSON-LD, structured-data examples, redirects, sitemaps, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for the complete historical token and record source IPs, ASNs, PTR results, forward lookups, paths, methods, response sizes, statuses, timing, rate, and redirect chains. Confirm whether the traffic is produced by a real legacy integration or a local test before attributing it to Google.
Use the current Rich Results Test for Google rich-result eligibility and Schema Markup Validator for generic Schema.org validation; do not build a new dependency on the retired endpoint. Ensure any robots group is exact and that private routes are protected by application authorization. Do not mistake noindex or robots directives for access control.
Re-check Google’s structured-data documentation and Search Central updates when tools or endpoints change. Keep this profile at legacy-label until a new source identifies a current client that should be documented separately. Do not claim successful blocking from configuration alone; verify subsequent logs and response behavior.
References
- Google structured-data testing guidance — current Google page recommending Rich Results Test and Schema Markup Validator and explaining the retired tool, reviewed with the headless browser on 2026-08-25.
- An update on the Structured Data Testing Tool — Google Search Central post describing deprecation and migration, reviewed with the headless browser on 2026-08-25.
- Rich Results Test — current Google-specific structured-data testing service.
- Schema Markup Validator — current generic Schema.org validation service.
- Google crawler verification guidance — use current reverse and forward DNS checks for claimed Google requests.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.