← Policy library/Apache Web Server: Blocking AI Crawlers & Scrapers
Policy library

Apache Web Server: Blocking AI Crawlers & Scrapers

Production-ready Apache HTTP Server (.htaccess and httpd.conf) configurations to block, throttle, or isolate aggressive AI crawlers and automated bots.

AI Summary: Web administrators can enforce deterministic AI bot blocking in Apache using mod_rewrite and mod_authz_core. Inspecting the HTTP_USER_AGENT header at the server level immediately returns HTTP 403 Forbidden without executing application code.

Enforcing Bot Policies at the Web Server Layer

While robots.txt establishes advisory crawling boundaries, rogue scrapers and high-concurrency training bots frequently ignore voluntary directives. Enforcing crawler policies inside Apache HTTP Server (httpd.conf, virtual hosts, or .htaccess) drops unauthorized connections before they reach your PHP, Node.js, or Python application runtimes.

Method 1: Using mod_rewrite (Recommended for .htaccess)

Ensure mod_rewrite is enabled in Apache. Place the following rules near the top of your document root .htaccess file:

configuration / code
<IfModule mod_rewrite.c>
    RewriteEngine On

    # Match known AI training crawlers and unauthorized scrapers
    RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|Bytespider|CCBot|Amazonbot|anthropic-ai|cohere-ai|Omgilibot) [NC]
    
    # Return immediate 403 Forbidden
    RewriteRule .* - [F,L]
</IfModule>

Method 2: Using mod_authz_core & SetEnvIf (Apache 2.4+)

For centralized virtual host configuration inside httpd.conf, SetEnvIfNoCase paired with Require not env offers superior throughput and lower memory overhead:

configuration / code
# Flag unauthorized AI user-agents
SetEnvIfNoCase User-Agent "GPTBot" block_ai_crawler
SetEnvIfNoCase User-Agent "ClaudeBot" block_ai_crawler
SetEnvIfNoCase User-Agent "Bytespider" block_ai_crawler
SetEnvIfNoCase User-Agent "CCBot" block_ai_crawler

<Directory "/var/www/html">
    <RequireAll>
        Require all granted
        Require not env block_ai_crawler
    </RequireAll>
</Directory>

Method 3: Protecting Specific Directories While Leaving Public Content Open

If your goal is to allow AI crawlers to index your public marketing pages while protecting sensitive documentation, customer portals, or download assets:

configuration / code
<Location "/internal-docs">
    <IfModule mod_rewrite.c>
        RewriteEngine On
        RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|PerplexityBot) [NC]
        RewriteRule .* - [F,L]
    </IfModule>
</Location>

Testing & Validation Commands

Always test your Apache configuration before and after applying rules:

configuration / code
# 1. Verify Apache configuration syntax
apache2ctl configtest

# 2. Test benign request (should return HTTP 200)
curl -I -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64)" https://example.com/

# 3. Test blocked AI crawler user-agent (must return HTTP 403)
curl -I -A "Mozilla/5.0 (compatible; GPTBot/1.0; +https://openai.com/gptbot)" https://example.com/

Need a comprehensive crawler policy that protects server resources while preserving AI search citations? Audit your site with Geolify.ai.

Related policies