Bot directory / search

Googlebot: Robots.txt & Crawl Policy Reference

Comprehensive guide to Googlebot, Google's primary web crawler. Learn how to manage search indexing and understand its crawl limits.

AI Summary: Googlebot is the generic name for Google's primary web crawler. It discovers and fetches web pages to build the Google Search index, serving as the most critical bot for general web visibility.

Role and policy boundary

Googlebot is responsible for crawling the web for Google Search. It operates continuously and renders JavaScript to understand modern web applications. Allowing Googlebot is essential for SEO. It is distinct from Google-Extended, which controls whether your content is used to improve Google's generative AI models.

A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting. Start with a dedicated group:

configuration / code
User-agent: Googlebot
Allow: /
Disallow: /staging/
Disallow: /internal/

To stop access for the entire site, use:

configuration / code
User-agent: Googlebot
Disallow: /

Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect assistant and training crawlers too, and it makes later audits harder because the source of the decision is less specific.

Layered verification

Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.

Googlebot strictly follows robots.txt rules and page-level X-Robots-Tag headers (like noindex). Notably, Googlebot has a documented file size limit: it only indexes the first 2 MB of an HTML file (or supported text-based file) for Search. Content beyond this limit is not processed.

The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why a bot is allowed while another is blocked, rather than returning one blended website score.

Page-level directives can still override the intended outcome for indexing:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

If a response uses these tags, the report marks the result as blocked or conflicting even if the crawler-specific robots group is permissive. This is especially important for canonical pages served through an edge cache where headers may differ from the origin response.

WAF and Nginx remediation examples

configuration / code
{
  "description": "Observe googlebot candidates",
  "expression": "lower(http.user_agent) contains \"googlebot\"",
  "action": "log"
}

To protect your site from bad actors spoofing Googlebot, rely on verified bot detection in your WAF or perform Reverse DNS verification:

configuration / code
# Conceptual Nginx configuration for verified bots
# Requires a robust IP allowlist or dynamic DNS resolution script
if ($http_user_agent ~* "Googlebot") {
    # Pass to backend if IP is verified
}

Use your platform's actual middleware response pattern rather than copying this simplified example without review. Never place a secret, verification token, or internal policy identifier in a public response header.

Review checklist

Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.

Verify that the dedicated group appears before relying on a wildcard, test a representative public and private path, and compare live response headers with robots.txt. Keep the policy close to the content owner's intent and record whether the site wants discovery, citation, or no access at all.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.