← Policy library/Training Opt-Out Policy: How to Block AI Data Scraping
Policy library

Training Opt-Out Policy: How to Block AI Data Scraping

A comprehensive guide to blocking AI crawlers from using your website's content for training large language models (LLMs) while maintaining search visibility.

AI Summary: A training opt-out policy allows website owners to declare that their content must not be used to train AI models. Because AI companies use different crawlers for search indexing and model training, webmasters can use robots.txt to block training bots (like GPTBot or Applebot-Extended) while still allowing search bots (like OAI-SearchBot or Applebot) to index their site for visibility.

The Separation of Search and Training

In the era of generative AI, content discovery and content extraction for model training are two distinct operations. Many major AI vendors now operate separate crawlers or provide specific robots.txt tokens to distinguish between these intents.

If you block a vendor's primary search crawler, your website may disappear from their search engine or AI-powered answer engine entirely. If you only block their training crawler, your site remains visible in search results, but the vendor is instructed not to scrape your content to train future models.

For example:

  • OpenAI uses OAI-SearchBot for SearchGPT visibility, but GPTBot for model training.
  • Apple uses Applebot for Siri and Spotlight search, but Applebot-Extended as an opt-out token for generative AI training.
  • Anthropic uses ClaudeBot for general web fetching and training.

Recommended Layering for Training Opt-Outs

To effectively opt out of AI training without sacrificing search visibility, you must explicitly declare rules for known training crawlers in your robots.txt file.

Step 1: Identify Training Crawlers

Maintain a list of known AI training crawlers. Common examples include:

  • GPTBot (OpenAI)
  • Applebot-Extended (Apple)
  • ClaudeBot (Anthropic)
  • Bytespider (ByteDance)
  • CCBot (Common Crawl)
  • FacebookBot (Meta)
  • Google-Extended (Google's token for Bard/Vertex AI training)

Step 2: Configure robots.txt

Add explicit Disallow rules for these crawlers. Ensure these rules are placed before any wildcard (*) rules.

configuration / code
# Block OpenAI's training crawler
User-agent: GPTBot
Disallow: /

# Block Apple's AI training token
User-agent: Applebot-Extended
Disallow: /

# Block Google's AI training token
User-agent: Google-Extended
Disallow: /

# Block Anthropic's crawler
User-agent: ClaudeBot
Disallow: /

# Block Common Crawl (often used by open-source models)
User-agent: CCBot
Disallow: /

# Allow general search engines
User-agent: *
Allow: /

Enforcement and WAF Remediation

While robots.txt is the standard method for declaring a training opt-out, it is purely advisory. Malicious or poorly-configured scrapers may ignore these directives.

To enforce your policy, you must implement network-level controls using a Web Application Firewall (WAF) or server configuration (like Nginx).

WAF JSON Example

You can create a WAF rule to block requests from known training User-Agents:

configuration / code
{
  "description": "Block known AI training crawlers",
  "expression": "lower(http.user_agent) contains \"gptbot\" or lower(http.user_agent) contains \"claudebot\" or lower(http.user_agent) contains \"bytespider\" or lower(http.user_agent) contains \"ccbot\"",
  "action": "block"
}

(Note: Applebot-Extended and Google-Extended are policy tokens, not User-Agents, so they will not appear in HTTP headers and cannot be blocked via WAF User-Agent rules. You must block the primary crawler, e.g., Applebot or Googlebot, if you wish to enforce the block at the network level, though this will also block search indexing.)

Nginx Remediation Example

You can block these agents directly in your web server configuration:

configuration / code
map $http_user_agent $block_ai_training {
    default 0;
    ~*GPTBot 1;
    ~*ClaudeBot 1;
    ~*Bytespider 1;
    ~*CCBot 1;
    ~*FacebookBot 1;
}

server {
    location / {
        if ($block_ai_training) {
            return 403;
        }
        # Standard configuration follows...
    }
}

Review Checklist

When implementing a training opt-out policy, verify the following:

  • [ ] Have you separated search crawlers (e.g., OAI-SearchBot) from training crawlers (e.g., GPTBot) in your robots.txt?
  • [ ] Are the training crawlers explicitly disallowed (Disallow: /)?
  • [ ] Have you implemented WAF or server-level blocks for crawlers that are known to ignore robots.txt or for which you require strict enforcement?
  • [ ] Have you tested your site's visibility on major search engines to ensure you haven't accidentally blocked general indexing?
  • [ ] Are you monitoring your server logs for new or spoofed AI crawler User-Agents?

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.

Related policies