← Bot Directory/Applebot-Extended
Bot directory / ai-training

Applebot-Extended: Robots.txt & Crawl Policy Reference

Technical reference for Applebot-Extended, Apple's dedicated robots.txt control token for generative AI model training.

AI Summary: Applebot-Extended is a secondary robots.txt control token provided by Apple. It is not an independent crawler but a policy directive that allows web publishers to opt out of having their content used for Apple's generative AI model training. Blocking Applebot-Extended prevents AI training use while still allowing the primary Applebot to index content for Siri and Spotlight search.

Role and policy boundary

According to Apple's official documentation, Applebot-Extended is a dedicated control token for generative AI model training. It does not perform independent crawling. Instead, Apple uses the primary Applebot crawler to fetch content, and then applies the Applebot-Extended rules to determine if that fetched content can be used to train its generative AI models.

If you want to allow Apple to index your site for Siri and Spotlight Suggestions but prohibit the use of your content for generative AI training, you must block Applebot-Extended.

To opt out of Apple's generative AI training across your entire site, add the following to your robots.txt:

configuration / code
User-agent: Applebot-Extended
Disallow: /

Because Applebot-Extended is only a policy token, blocking it does not reduce crawl load. The primary Applebot will continue to crawl your site according to its own rules. To block both search indexing and AI training, you must block Applebot:

configuration / code
User-agent: Applebot
Disallow: /

These are explanatory site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.

Layered verification

Because Applebot-Extended is a policy token and not a distinct crawler, you will not see Applebot-Extended in your server access logs. The actual requests are made by Applebot.

To verify Applebot traffic, start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate.

Apple officially supports reverse-DNS verification. To authenticate an Applebot request:

  1. Perform a reverse DNS lookup on the accessing IP address.
  2. Verify that the hostname ends with .applebot.apple.com.
  3. Perform a forward DNS lookup on that hostname to confirm it resolves back to the original IP address.

Do not treat the word Applebot in a header as proof of ownership. Check source IP and DNS evidence independently before creating a trust exception.

Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish.

Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate a crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

Since Applebot-Extended does not appear in HTTP headers, WAF and Nginx rules must target the primary Applebot if you wish to enforce access controls at the network edge. Use report-only logging and a narrow observation first:

configuration / code
{
  "description": "Observe unverified Applebot candidates",
  "expression": "lower(http.user_agent) contains \"applebot\"",
  "action": "log"
}

After independent DNS verification and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule:

configuration / code
map $http_user_agent $block_applebot_private {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*Applebot 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_applebot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Remember that Applebot-Extended is a robots.txt directive, not a User-Agent you will find in logs. Search logs for Applebot instead, and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Re-check the crawler registry, Apple's official documentation, robots.txt, and operator contact path. Verify Applebot IPs using the .applebot.apple.com reverse-DNS method.

Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.

Review search indexing, AI input, reference use, and model-training decisions separately. Apple explicitly separates search indexing (Applebot) from AI training (Applebot-Extended). Do not claim successful blocking or verification from configuration alone; validate later logs.

Record the date and reason for the documented classification so future evidence can be compared. If a new deployment identifies itself differently, create a separate evidence trail.

References

  1. About Applebot — official documentation detailing Applebot, Applebot-Extended, and reverse-DNS verification.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.