← Bot Directory/DuckDuckGo-Favicons-Bot
Bot directory / search-engine

DuckDuckGo-Favicons-Bot: Robots.txt & Crawl Policy Reference

Technical reference for DuckDuckGo-Favicons-Bot, the official favicon crawler for DuckDuckGo. Learn how to verify its traffic and manage its access.

AI Summary: DuckDuckGo-Favicons-Bot is an official web crawler operated by DuckDuckGo, specifically designed to fetch site icons (favicons) for display in their search results and browser products. It respects standard robots.txt directives. Blocking this bot will prevent your site's favicon from appearing next to your listings in DuckDuckGo search results. Verify its traffic using reverse DNS lookups to ensure the IPs belong to *.duckduckgo.com.

Role and policy boundary

DuckDuckGo-Favicons-Bot is a specialized crawler for the DuckDuckGo search engine. Its sole purpose is to discover and retrieve the favicon associated with a website, enhancing the visual presentation of search results. The crawler identifies itself using variations of the User-Agent string, most commonly containing DuckDuckGo-Favicons-Bot (e.g., Mozilla/5.0 (compatible; DuckDuckGo-Favicons-Bot/1.0; +http://duckduckgo.com)).

As an official search engine crawler, its primary purpose is public indexing of visual assets. Blocking DuckDuckGo-Favicons-Bot means your site will appear with a generic or missing icon in DuckDuckGo search results. Do not infer that its primary purpose is AI training.

If your logs confirm an exact DuckDuckGo-Favicons-Bot token and you want to prevent your favicon from being fetched (though generally not recommended for branding purposes), publish:

configuration / code
User-agent: DuckDuckGo-Favicons-Bot
Disallow: /

For selective access (e.g., allowing it only to access the root directory where favicons usually reside):

configuration / code
User-agent: DuckDuckGo-Favicons-Bot
Allow: /favicon.ico
Allow: /apple-touch-icon.png
Disallow: /private/
Disallow: /internal/
Disallow: /api/

Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Because it is a search engine crawler, it may be spoofed by malicious scrapers, so you must verify its authenticity.

DuckDuckGo provides a verification method via reverse DNS. Perform a reverse DNS (PTR) lookup on the accessing IP address. A genuine DuckDuckGo-Favicons-Bot IP will resolve to a hostname ending in *.duckduckgo.com. You can then perform a forward DNS lookup on that hostname to ensure it matches the original IP.

Analyze behavior without assigning purpose prematurely. Requests for image files like favicon.ico or apple-touch-icon.png in the root directory resemble authorized favicon indexing; deep traversal of private areas, high concurrency ignoring crawl delays, repeated retries, or private API access may indicate spoofing. These patterns demonstrate operational impact but cannot prove the operator or downstream use.

Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:

configuration / code
<meta name="DuckDuckGo-Favicons-Bot" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.

WAF and Nginx remediation examples

Once logs confirm an exact unwanted token (e.g., a spoofed bot failing reverse DNS verification), a narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:

configuration / code
{
  "description": "Block observed DuckDuckGo-Favicons-Bot token",
  "expression": "lower(http.user_agent) contains \"duckduckgo-favicons-bot\"",
  "action": "block"
}

For Nginx, scope enforcement to private and high-cost routes while investigating public access:

configuration / code
map $http_user_agent $block_duckduckgo_favicons_bot {
    default 0;
    ~*DuckDuckGo-Favicons-Bot 1;
}

server {
    location ~ ^/(private|internal|account|uploads|paywall|api)/ {
        if ($block_duckduckgo_favicons_bot) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and may block a legitimate client using the substring. Do not create an IP allowlist or broad network block without current operator evidence. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.

Review checklist

Search logs for every exact header that may be associated with DuckDuckGo-Favicons-Bot and preserve representative requests. Record paths, response sizes, statuses, source networks, timing, and rate. Verify the traffic by performing a reverse DNS lookup to ensure the hostname ends in *.duckduckgo.com.

Decide whether your objective is to preserve visual branding on DuckDuckGo, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group only for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately.

References

  1. DuckDuckBot Official Documentation — official documentation portal for DuckDuckGo crawlers.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.