← Bot Directory/APIs-Google
Bot directory / scraper

APIs-Google: Robots.txt & Crawl Policy Reference

Technical reference for APIs-Google, a special-case crawler by Google. Learn how to verify its traffic and manage its access, noting that it ignores global robots.txt rules.

AI Summary: APIs-Google is a "special-case" crawler operated by Google, typically used by specific Google APIs and products to fetch web content. Crucially, Google's documentation states that APIs-Google ignores global User-agent: * rules in robots.txt. To control its access, you must target it explicitly. It operates from a distinct IP range verifiable via reverse DNS (rate-limited-proxy-***.google.com).

Role and policy boundary

According to Google's official crawler documentation, APIs-Google is classified as a special-case crawler. It identifies itself in HTTP requests as APIs-Google (+https://developers.google.com/webmasters/APIs-Google.html). It is used by various Google services and APIs to retrieve content on behalf of those services.

A critical policy boundary for APIs-Google is its interaction with robots.txt. Unlike standard crawlers (like Googlebot), APIs-Google ignores wildcard directives. If you rely on User-agent: * to block traffic, APIs-Google will bypass those restrictions.

To communicate a restriction to this specific crawler, you must publish an explicit rule:

configuration / code
User-agent: APIs-Google
Disallow: /

For selective access:

configuration / code
User-agent: APIs-Google
Allow: /public-api-data/
Disallow: /private/
Disallow: /internal/

Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate.

Google provides robust verification for its crawlers. For special-case crawlers like APIs-Google, the IP ranges are published in Google's special-crawlers.json object. Furthermore, you can verify the traffic using reverse DNS lookups. The reverse DNS mask for these special-case crawlers matches rate-limited-proxy-***-***-***-***.google.com.

Analyze behavior without assigning purpose prematurely. Requests for public resources may be legitimate API fetches; deep traversal, high concurrency, repeated retries, original-asset downloads, or private API access may indicate misconfiguration or spoofing by actors pretending to be Google.

Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.

WAF and Nginx remediation examples

Once logs confirm an exact unwanted token (e.g., a spoofed bot failing reverse DNS verification, or if you explicitly want to block Google's API fetchers), a narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:

configuration / code
{
  "description": "Block observed APIs-Google token",
  "expression": "lower(http.user_agent) contains \"apis-google\"",
  "action": "block"
}

For Nginx, scope enforcement to private and high-cost routes while investigating public access:

configuration / code
map $http_user_agent $block_apis_google {
    default 0;
    ~*APIs-Google 1;
}

server {
    location ~ ^/(private|internal|account|uploads|paywall)/ {
        if ($block_apis_google) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and may block a legitimate client using the substring. Do not create an IP allowlist or broad network block without current operator evidence. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.

Review checklist

Search logs for every exact header that may be associated with APIs-Google and preserve representative requests. Record paths, response sizes, statuses, source networks, timing, and rate. Verify the traffic against Google's special-crawlers.json or by performing a reverse DNS lookup ensuring the hostname ends in google.com.

Decide whether your objective is to preserve API integration, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group only for the exact observed token (remembering that global rules are ignored), enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately.

References

  1. Google Special-Case Crawlers Documentation — official documentation detailing APIs-Google's behavior and robots.txt handling, accessed 2026-08-24.
  2. Google Verifying Crawlers — official guide on verifying Googlebot and other Google crawlers via IP and DNS.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.