← Bot Directory/Google Favicon
Bot directory / search-engine

Google Favicon: Robots.txt & Crawl Policy Reference

Technical reference for the community-recorded Google Favicon User-Agent pattern, with explicit limits around Google ownership, IP verification, and current policy.

AI Summary: Google Favicon is a community-recorded User-Agent pattern in the crawler-user-agents project. The reviewed repository is not Google-owned and no Google first-party source verified this exact header, its current activity, IP ranges, or a dedicated policy. Treat it as an unverified favicon-discovery label and verify requests from logs before applying controls.

Role and policy boundary

The registry entry associates this label with favicon retrieval and records the complete pattern:

configuration / code
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/49.0.2623.75 Safari/537.36 Google Favicon

The linked repository describes itself as a community-maintained list of syntactic User-Agent patterns used by robots, crawlers, and spiders. It is useful for detection, but it is not a Google policy page, an infrastructure allowlist, or proof that a matching request is sent by Google. The reviewed JSON did not independently establish a dedicated Google Favicon record with a first-party URL, and no Google-owned source reviewed in this pass confirmed this exact suffix.

Treat favicon discovery as an unverified, narrow purpose hypothesis. A request for /favicon.ico, manifest icons, or other site branding assets may be consistent with favicon retrieval, but that path pattern does not identify the operator. Do not infer Google Search indexing, AI-training use, current activity, Google infrastructure, or permission to access private paths from the name or User-Agent.

If your logs confirm this exact token and you want to exclude it from discovery, a site-owner robots rule could be:

configuration / code
User-agent: Google Favicon
Disallow: /

For a selective asset policy, approve only the public branding paths you intend to serve:

configuration / code
User-agent: Google Favicon
Allow: /favicon.ico
Allow: /icons/
Allow: /assets/brand/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

These are defensive examples, not published Google instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No Google first-party source reviewed here published a Google Favicon IP range, DNS verification procedure, or rate limit. Do not extend Googlebot verification rules to this label without a current Google source that expressly covers it.

Compare the observed path with the favicon hypothesis without treating it as attribution. A small request for a public favicon may be consistent with the stated label; a broad crawl, private endpoint access, repeated retries, high concurrency, or traffic from an unverified network is not proof of an official Google client. These observations establish operational impact, not operator identity or downstream use.

Evaluate /robots.txt independently. Confirm that the response is served from the intended host, has a successful status and text content type, and contains the exact group you choose to publish. Test the full header, not only a generic Google substring, and check whether a global group changes the intended result. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives do not authenticate a Google service or secure private routes. They also do not prove that a community-listed client has accepted the preference. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

If logs show a repeatable but unverified token, begin in report-only mode. Adapt the expression to your WAF provider and label the event as a pattern match rather than an official Google identity:

configuration / code
{
  "description": "Review observed Google Favicon pattern",
  "expression": "lower(http.user_agent) contains \"google favicon\"",
  "action": "log"
}

After reviewing false positives and confirming the business decision, scope an enforcement rule to asset or sensitive routes rather than broadly blocking all Google traffic:

configuration / code
map $http_user_agent $review_google_favicon {
    default 0;
    ~*Google[ ]Favicon 1;
}

server {
    location ~ ^/(admin|account|private|internal|api)/ {
        if ($review_google_favicon) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Do not add an IP allowlist, reverse-DNS suffix, or Googlebot exception based only on the community repository. A User-Agent is easy to spoof and a broad Google match may break legitimate search, Ads, or monitoring traffic. Test favicon files, manifests, public HTML, sitemaps, media, account flows, APIs, and approved integrations separately. Pair edge rules with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search logs for the exact registered suffix and record representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Confirm whether requests are limited to public icon assets and whether the source can be verified through an independent first-party document. Do not classify a request as Google traffic from a syntactic match alone.

Re-check the community registry for changes, then look for a current Google-owned crawler documentation page before upgrading this profile. Keep it at partially-documented until Google publishes a current canonical User-Agent, dedicated purpose, source-verification method, or applicable robots guidance. If a future request uses a different Google crawler token, record it against that token rather than broadening this one.

Decide whether the objective is to preserve favicon discovery, limit asset requests, protect private material, or reduce load. Publish exact robots rules and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent logs, cache behavior, and legitimate Google traffic separately.

References

  1. crawler-user-agents repository — community-maintained User-Agent pattern list linked by the registry, reviewed 2026-08-25.
  2. crawler-user-agents JSON data — raw community data reviewed for Google-related records; it does not authenticate Google infrastructure.
  3. Google crawler overview — Google’s first-party overview for documented Google crawlers; the reviewed material did not establish this exact Google Favicon pattern.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.