Bot directory / search-engine

Slurp: Robots.txt & Crawl Policy Reference

Technical reference for Yahoo! Slurp, the official Yahoo Search crawler, including its published User-Agent, indexing purpose, and robots.txt compliance.

AI Summary: Slurp is the official web crawler for Yahoo Search. Yahoo’s documentation publishes the exact User-Agent Mozilla/5.0 (compatible; Yahoo! Slurp; http://help.yahoo.com/help/us/ysearch/slurp) and confirms it obeys the Slurp robots.txt group. It indexes content for Yahoo Mobile Search and collects data for Yahoo News, Finance, and Sports. The registry records a historical /3.0 version, but the current official documentation omits the version number. Verify the source network before trusting the header.

Role and policy boundary

Yahoo’s official Help Central documents Slurp as the Yahoo Search robot for crawling and indexing web page information. While some Yahoo Search results are powered by partners (like Bing), Yahoo maintains Slurp to populate Yahoo Mobile Search and to collect content for partner sites like Yahoo News, Yahoo Finance, and Yahoo Sports.

The official documentation publishes this exact User-Agent:

configuration / code
Mozilla/5.0 (compatible; Yahoo! Slurp; http://help.yahoo.com/help/us/ysearch/slurp)

The registry records Yahoo! Slurp/3.0, but the current published string does not include the version number. Expect variations in the version suffix, but rely on the core Yahoo! Slurp identifier.

Yahoo explicitly states that webmasters can prevent Slurp from reading pages by using robots.txt. For selective access to public pages:

configuration / code
User-agent: Slurp
Allow: /public/
Allow: /news/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

For a complete exclusion:

configuration / code
User-agent: Slurp
Disallow: /

These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. The search-indexing purpose does not establish permission for third-party model training or generative AI input beyond Yahoo's specified services.

Layered verification

Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs. The reviewed Yahoo documentation does not publish a static IP list or a specific reverse-DNS verification procedure for Slurp.

Because a User-Agent is easily spoofed, do not treat the header alone as authentication. Use standard network verification: perform a reverse DNS lookup on the source IP and look for a yahoo.com or yahoo.net hostname, then perform a forward lookup to confirm it resolves back to the same IP. A Yahoo! Slurp header from an unverified IP or a residential network should not be treated as authentic Yahoo traffic.

Compare observed traffic with the documented search-indexing and news-gathering roles. Public HTML, metadata, links, feeds, sitemaps, and ordinary assets are consistent with discovery. Private endpoints, account routes, APIs, licensed content, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not permission.

Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact Slurp group you intend to publish. Check precedence, path matching, and actual request behavior. Meta directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

Begin in report-only mode and correlate the Yahoo! Slurp User-Agent with reverse/forward DNS verification. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe Yahoo Slurp candidates before enforcement",
  "expression": "lower(http.user_agent) contains \"yahoo! slurp\"",
  "action": "log"
}

For a deliberately restricted private route, use the exact token only after verifying the source network:

configuration / code
map $http_user_agent $block_slurp_private {
    default 0;
    ~*Yahoo!\ Slurp 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_slurp_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

This rule is route-scoped and does not prove identity. Do not invent an IP allowlist, Yahoo ASN exception, rate, or training permission without operator documentation.

A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact Yahoo! Slurp string, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Perform reverse and forward DNS verification to confirm a Yahoo hostname before trusting the client. Treat requests from unverified networks as spoofed.

Check your robots file for an exact Slurp group. Test precedence and path behavior, and verify actual logs after publishing. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.

Review search indexing, AI input, reference use, and model-training decisions separately. Yahoo explicitly states the crawler is used for Yahoo Mobile Search and Yahoo content properties.

Keep the profile at documented: the operator publishes the exact User-Agent, robots identifier, and search purpose. Re-check the crawler documentation when the source network or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.

References

  1. Why is Slurp crawling my page? — official Yahoo Help page documenting the User-Agent, robots compliance, and indexing purpose; reviewed with the headless browser on 2026-08-25.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.