← Bot Directory/Googlebot-News
Bot directory / search-engine

Googlebot-News: Robots.txt & Crawl Policy Reference

Technical reference for Googlebot-News, Google's dedicated crawler for indexing news content for Google News.

AI Summary: Googlebot-News is the official web crawler used by Google specifically to discover and index content for Google News. It crawls news publishers more frequently to capture time-sensitive articles. It fully respects robots.txt directives targeted at the Googlebot-News token. Google provides robust verification methods, including a published IP list and reverse-DNS lookups, allowing site owners to authenticate traffic safely.

Role and policy boundary

Google operates Googlebot-News to crawl the web specifically for Google News. This crawler targets news publishers and blogs, fetching content rapidly to ensure breaking news appears in Google's news surfaces in a timely manner.

According to Google's official Publisher Center documentation, Googlebot-News respects robots.txt. If you want to allow Google to index your site for general search but prevent your articles from appearing in Google News, you can block Googlebot-News while allowing Googlebot.

To exclude your content from Google News across your entire site, add the following to your robots.txt:

configuration / code
User-agent: Googlebot-News
Disallow: /

For selective access (e.g., preventing the indexing of specific non-news directories from Google News):

configuration / code
User-agent: Googlebot-News
Disallow: /press-releases/
Disallow: /internal-updates/

These are explanatory site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate.

Google officially supports two methods to verify Googlebot-News traffic:

  1. IP Allowlist: Google publishes a JSON file containing all active Googlebot IP ranges.
  2. Reverse DNS: Perform a reverse DNS lookup on the accessing IP address. Verify that the hostname ends with .googlebot.com or .google.com. Then, perform a forward DNS lookup on that hostname to confirm it resolves back to the original IP address.

Do not treat the word Googlebot-News in a header as proof of ownership. Malicious actors frequently spoof Googlebot User-Agents to bypass security controls. Check source IP and DNS evidence independently before creating a trust exception.

Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish.

Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate a crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

When configuring WAF rules, always pair User-Agent matching with IP verification. Avoid broad googlebot rules that blindly trust the header:

configuration / code
{
  "description": "Observe unverified Googlebot-News candidates",
  "expression": "lower(http.user_agent) contains \"googlebot-news\"",
  "action": "log"
}

After independent verification (using Google's IP list or reverse-DNS) and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule:

configuration / code
map $http_user_agent $is_googlebot_news {
    default 0;
    ~*Googlebot-News 1;
}

# Note: In a real Nginx configuration, you must pair this with IP validation
# (e.g., using the geo module with Google's published IP list) to prevent spoofing.
server {
    location ~ ^/(admin|account|private|licensed|internal)/ {
        # Only allow verified Google IPs here, or block entirely
        if ($is_googlebot_news) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact Googlebot-News substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Re-check the crawler registry, Google's official documentation, robots.txt, and operator contact path. Verify the IPs using Google's published list or the .googlebot.com reverse-DNS method.

Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.

Review search indexing, AI input, reference use, and model-training decisions separately. Googlebot-News is explicitly for Google News indexing. Do not claim successful blocking or verification from configuration alone; validate later logs.

Record the date and reason for the documented classification so future evidence can be compared. If a new deployment identifies itself differently, create a separate evidence trail.

References

  1. Google Publisher Center: Block access to content — official documentation on blocking Googlebot-News.
  2. Verifying Googlebot — instructions for verifying Google crawlers via IP or DNS.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.