← Bot Directory/Google-Site-Verification
Bot directory / search-engine

Google-Site-Verification: Robots.txt & Crawl Policy Reference

Technical reference for Google's Site Verifier fetcher, including ownership-verification methods, redirect behavior, robots considerations, and layered request validation.

AI Summary: Google documents the Google Site Verifier user agent as part of Search Console site-ownership verification. Verification may use an HTML file, an HTML tag, Analytics, Tag Manager, or other supported methods; tag-based checks follow the page reached by a non-logged-in redirect, while file-upload verification does not follow redirects. Treat it as user-triggered verification traffic, verify the source, and protect private routes with application controls rather than User-Agent rules alone.

Role and policy boundary

Google’s Search Console Help documentation states that Google uses the Google Site Verifier user agent to perform site ownership verification. The registry records this exact form:

configuration / code
Mozilla/5.0 (compatible; Google-Site-Verification/1.0)

The documented role is ownership verification, not general search indexing. Google lists HTML file upload, HTML tag, Google Analytics tracking code, Google Tag Manager container snippet, Google Sites, Blogger, and domain-name-provider methods. A verifier may fetch a public file, a redirected page, or a page containing a token so that Search Console can confirm control of the property.

Verification is a product function and may be user-triggered. A matching header does not grant access to private content, bypass authentication, or prove that the requester should be allowed through a paywall. Do not infer AI-training use, broad search crawling, or permission to access arbitrary paths from the verifier identity.

If you need a robots rule for verification traffic, first confirm which verification method your site uses and whether restricting the path would prevent ownership checks. A selective example is:

configuration / code
User-agent: Google-Site-Verification
Allow: /google-site-verification.html
Allow: /site-verification/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

For a deliberately restricted verifier, use an exact group:

configuration / code
User-agent: Google-Site-Verification
Disallow: /

The second rule can make verification fail and should not be added without understanding the ownership workflow. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Begin with access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS hostname, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Google’s crawler overview says Google clients identify themselves through the HTTP User-Agent, source IP, and reverse DNS hostname. Use all three signals when attributing a request.

Follow Google’s current verification guidance for a claimed source: reverse-resolve the requesting IP, assess the hostname against Google’s documented naming patterns, then forward-resolve the hostname to confirm it maps back to the original address. Keep ownership verification separate from ordinary Googlebot crawling. Do not classify an arbitrary request as authentic from the Google-Site-Verification string alone.

Validate the verification workflow itself. For tag-based methods, Google says Search Console looks for the verification tag on the page to which a non-logged-in visitor is redirected from the property URL. For file upload, redirects are not followed. Check the exact property URL, redirect chain, final response status, body, content type, and presence of the expected token without exposing private credentials.

Evaluate /robots.txt independently. Confirm that the file is served by the correct host, returns a successful status and text content type, and allows the exact file or page used by your chosen verification method. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives do not authenticate a verifier or secure private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

When troubleshooting a failed verification, use a report-only rule first. Adapt the expression to your WAF provider and correlate it with the verification path and independently checked source:

configuration / code
{
  "description": "Observe Google Site Verification requests",
  "expression": "lower(http.user_agent) contains \"google-site-verification\"",
  "action": "log"
}

For Nginx, preserve the public verification file while restricting sensitive routes:

configuration / code
map $http_user_agent $block_google_site_verifier {
    default 0;
    ~*Google-Site-Verification 1;
}

server {
    location = /google-site-verification.html {
        try_files $uri =404;
    }

    location ~ ^/(admin|account|private|internal|api)/ {
        if ($block_google_site_verifier) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and can break a legitimate ownership check if applied to the verification file. Do not broadly block Google clients or create a stale IP allowlist. Test HTML file upload, tag-based verification, redirected URLs, analytics/tag-manager methods, public HTML, sitemaps, media, accounts, and APIs separately. Pair edge rules with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Identify the verification method before changing robots or WAF rules. Search logs for the exact registered token and record source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, and redirect chain. Confirm that the expected verification file or tag is reachable by the required anonymous request.

Check that robots.txt permits the chosen public verification path and that no middleware, WAF, authentication wall, cache rule, or redirect changes the response. Remember that Search Console’s tag-based check follows the page reached by a non-logged-in redirect, while file-upload verification does not follow redirects.

If you restrict the verifier, verify that the intended Search Console ownership check fails or succeeds for the reason you expect, and confirm that Googlebot and other intentionally allowed identities remain unaffected. Re-check Google’s documentation when verification methods or User-Agent strings change.

Do not claim successful blocking or successful verification from a configuration change alone. Verify subsequent access logs, response body, headers, status, and Search Console result through an authorized site-owner workflow.

References

  1. Verify your site ownership — Google Search Console documentation describing the Google Site Verifier user agent and ownership-verification methods, reviewed with the headless browser on 2026-08-25.
  2. Google crawlers and fetchers overview — Google documentation for crawler categories and User-Agent, source IP, and reverse-DNS verification.
  3. Google crawler verification guidance — use current reverse and forward DNS checks for claimed Google requests.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.