← Bot Directory/anthropic-ai
Bot directory / ai-training

How to block anthropic-ai: Anthropic's AI Training Crawler

A complete technical reference for anthropic-ai, the crawler used by Anthropic to collect data for training the Claude family of AI models.

AI Summary: The anthropic-ai crawler is operated by Anthropic specifically to gather data for training their Claude AI models. To prevent your content from being used in Claude's training datasets, you must block User-agent: anthropic-ai in your robots.txt or block its User-Agent string at your firewall.

Role and policy boundary

Anthropic operates multiple crawlers with distinct purposes. The anthropic-ai crawler is dedicated entirely to data collection for training their foundation models, such as Claude 3. This is distinct from ClaudeBot, which is primarily used for real-time data retrieval when a user prompts the Claude assistant with a specific URL.

Understanding this boundary is crucial for data governance. If you wish to allow Claude users to summarize or interact with your live pages, you may choose to allow ClaudeBot. However, if you simultaneously want to protect your intellectual property from being permanently ingested into Anthropic's training corpus, you must explicitly block anthropic-ai.

Layered verification

To effectively block anthropic-ai, you should implement a multi-layered defense strategy.

The first layer is the robots.txt file, which Anthropic respects for training data collection. By specifying the anthropic-ai token, you signal your intent to opt out of their training datasets.

For a more robust defense, especially against aggressive crawling or if the robots.txt is bypassed, network-level blocking is highly recommended. The anthropic-ai crawler identifies itself using a specific User-Agent string (e.g., Mozilla/5.0 (compatible; anthropic-ai/1.0; +https://www.anthropic.com/ai)). By configuring your WAF or web server to drop requests matching this string, you create a hard barrier that enforces your data policy regardless of robots.txt parsing.

WAF and Nginx remediation examples

To block the anthropic-ai crawler at the server level, you can use User-Agent matching rules.

For an Nginx server, add the following configuration to return a 403 Forbidden status for requests originating from this crawler:

configuration / code
if ($http_user_agent ~* (anthropic-ai)) {
    return 403;
}

If your infrastructure relies on Cloudflare WAF, you can implement a custom rule using the following JSON expression to block the traffic at the edge:

configuration / code
{
  "action": "block",
  "expression": "(http.user_agent contains \"anthropic-ai\")",
  "description": "Block Anthropic training crawler (anthropic-ai)"
}

Review checklist

Ensure your opt-out policy for Anthropic's training data is effective by reviewing these steps:

  1. Confirm that User-agent: anthropic-ai is listed in your robots.txt with a Disallow: / directive.
  2. Verify that your WAF or server configuration is actively blocking the anthropic-ai User-Agent string.
  3. Differentiate your policy for anthropic-ai (training) and ClaudeBot (user-prompted retrieval) based on your business needs.
  4. Test the configuration by simulating a request with the anthropic-ai User-Agent to ensure it is rejected.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.