← Bot Directory/Googlebot-Mobile
Bot directory / search-engine

Googlebot-Mobile: Robots.txt & Crawl Policy Reference

Technical reference for Googlebot-Mobile, a legacy crawler token used by Google for mobile search indexing.

AI Summary: Googlebot-Mobile is a historical web crawler token used by Google specifically for indexing mobile web content (such as WAP or early feature phone sites). Today, Google primarily uses the standard Googlebot token with a smartphone user-agent string for mobile-first indexing. While Googlebot-Mobile is largely legacy, it is an official Google token that respects robots.txt and can be verified using Google's published IP list and reverse-DNS methods.

Role and policy boundary

Google previously operated Googlebot-Mobile to crawl content intended specifically for feature phones (like i-mode, DoCoMo, etc.). The registry records historical User-Agents such as DoCoMo/2.0 N905i... (compatible; Googlebot-Mobile/2.1; +http://www.google.com/bot.html).

With the shift to mobile-first indexing, Google now predominantly uses the standard Googlebot token accompanied by a modern smartphone User-Agent string (e.g., simulating an Android device). However, Googlebot-Mobile remains a documented part of Google's crawling history and infrastructure. It respects robots.txt directives.

If you have legacy rules or specific needs to block this older token, you can add the following to your robots.txt:

configuration / code
User-agent: Googlebot-Mobile
Disallow: /

Note that blocking Googlebot-Mobile will not prevent modern mobile-first indexing by Google. To manage how Google crawls your site today, you should target the primary Googlebot token:

configuration / code
User-agent: Googlebot
Disallow: /

These are explanatory site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate.

Google officially supports two methods to verify Googlebot traffic:

  1. IP Allowlist: Google publishes a JSON file containing all active Googlebot IP ranges.
  2. Reverse DNS: Perform a reverse DNS lookup on the accessing IP address. Verify that the hostname ends with .googlebot.com or .google.com. Then, perform a forward DNS lookup on that hostname to confirm it resolves back to the original IP address.

Do not treat the word Googlebot-Mobile in a header as proof of ownership. Malicious actors frequently spoof Googlebot User-Agents to bypass security controls. Check source IP and DNS evidence independently before creating a trust exception.

Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish.

Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate a crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

When configuring WAF rules, always pair User-Agent matching with IP verification. Avoid broad googlebot rules that blindly trust the header:

configuration / code
{
  "description": "Observe unverified Googlebot-Mobile candidates",
  "expression": "lower(http.user_agent) contains \"googlebot-mobile\"",
  "action": "log"
}

After independent verification (using Google's IP list or reverse-DNS) and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule:

configuration / code
map $http_user_agent $is_googlebot_mobile {
    default 0;
    ~*Googlebot-Mobile 1;
}

# Note: In a real Nginx configuration, you must pair this with IP validation
# (e.g., using the geo module with Google's published IP list) to prevent spoofing.
server {
    location ~ ^/(admin|account|private|licensed|internal)/ {
        # Only allow verified Google IPs here, or block entirely
        if ($is_googlebot_mobile) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact Googlebot-Mobile substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Re-check the crawler registry, Google's official documentation, robots.txt, and operator contact path. Verify the IPs using Google's published list or the .googlebot.com reverse-DNS method.

Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.

Review search indexing, AI input, reference use, and model-training decisions separately. Do not claim successful blocking or verification from configuration alone; validate later logs.

Record the date and reason for the documented classification so future evidence can be compared. If a new deployment identifies itself differently, create a separate evidence trail.

References

  1. Google Search Central: Googlebot — official documentation on Googlebot and its variants.
  2. Verifying Googlebot — instructions for verifying Google crawlers via IP or DNS.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.