← Policy library/Robots Layering Policy: Coordinating Access Controls
Policy library

Robots Layering Policy: Coordinating Access Controls

A technical guide to understanding how robots.txt, page-level metadata, HTTP headers, and WAF rules interact to control AI bot access.

AI Summary: Controlling AI bots requires a multi-layered approach. robots.txt signals crawling intent, HTML meta tags and HTTP headers signal indexing intent, and server/WAF configurations enforce security. A block in any of these layers should be explicitly documented. Webmasters must understand how these layers interact to effectively manage AI search visibility, model training opt-outs, and data security.

The Hierarchy of Bot Controls

Managing how AI crawlers interact with your website cannot be achieved with a single tool. It requires a coordinated strategy across four distinct layers of your infrastructure.

  1. Crawling Intent (robots.txt): The first line of communication. It tells bots which paths they are allowed to request.
  2. Indexing Intent (Meta Tags & Headers): Instructions on what the bot should do with the content after it has fetched it (e.g., do not index, do not follow links).
  3. Application Security (Authentication/Authorization): The logic that ensures private data is only accessible to verified users.
  4. Network Enforcement (WAF/CDN/Edge): The hard barrier that actively blocks malicious, spoofed, or non-compliant bots based on IP, behavior, or headers.

A critical rule of robots layering: A permissive robots.txt cannot make a protected route public, but a restrictive server rule will override a permissive robots.txt.

Recommended Layering Strategy

To build a robust defense and optimization strategy for AI bots, implement controls at each layer.

Layer 1: Crawling Intent (robots.txt)

Use robots.txt to define broad access rules based on bot categories (Search, Training, Assistant).

configuration / code
# 1. Allow AI Search bots for visibility
User-agent: OAI-SearchBot
Allow: /
Disallow: /private/

# 2. Allow AI Assistants to fetch pages for users
User-agent: ChatGPT-User
Allow: /
Disallow: /private/

# 3. Block AI Training bots to protect intellectual property
User-agent: GPTBot
User-agent: CCBot
Disallow: /

Layer 2: Indexing Intent (Page-Level & Headers)

If a bot is allowed to crawl a page, you can still control how it processes the data using HTML meta tags or the X-Robots-Tag HTTP header.

HTML Meta Tag:

configuration / code
<!-- Prevent the page from being indexed in search results -->
<meta name="robots" content="noindex, nofollow">

HTTP Header (Useful for non-HTML files like PDFs or Images):

configuration / code
X-Robots-Tag: noindex, noarchive

Note: For a bot to see a noindex tag, it must be allowed to crawl the page in robots.txt. If you block a page in robots.txt, the bot will never see the noindex tag, and the URL might still appear in search results (though without a description).

Layer 3: Application Security

Never rely on robots.txt or meta tags to protect sensitive data. AI bots do not execute JavaScript in the same way browsers do, and they do not possess user session cookies unless explicitly provided.

Ensure your backend application returns a 401 Unauthorized or 403 Forbidden for any request lacking a valid session token, regardless of the User-Agent.

Layer 4: Network Enforcement (WAF & Nginx)

Use your Web Application Firewall (WAF) or reverse proxy (like Nginx) to enforce your policies against bots that ignore robots.txt or attempt to spoof their identities.

WAF and Nginx Remediation Examples

WAF JSON Example

Enforce a hard block on known training bots at the edge, ensuring they consume zero server resources:

configuration / code
{
  "description": "Hard block on AI training bots at the edge",
  "expression": "lower(http.user_agent) contains \"gptbot\" or lower(http.user_agent) contains \"ccbot\"",
  "action": "block"
}

Nginx Remediation Example

Coordinate your Nginx configuration to respect the layers: allow public access, enforce authentication for private routes, and apply specific blocks based on bot identity.

configuration / code
map $http_user_agent $is_training_bot {
    default 0;
    ~*GPTBot 1;
    ~*CCBot 1;
}

server {
    # Layer 2: Add X-Robots-Tag to sensitive file types globally
    location ~* \.(pdf|docx|zip)$ {
        add_header X-Robots-Tag "noindex, nofollow";
        try_files $uri =404;
    }

    # Layer 4: Hard block training bots from the entire site
    if ($is_training_bot) {
        return 403;
    }

    # Layer 3: Enforce Application Security on private routes
    location ~ ^/(admin|private|billing)/ {
        # Real authentication required
        auth_request /auth-verify;
        error_page 401 = /login;

        # Even if auth succeeds, add noindex as a failsafe
        add_header X-Robots-Tag "noindex";
    }

    # Public routes
    location / {
        try_files $uri $uri/ /index.html;
    }
}

Review Checklist

When auditing your robots layering strategy, verify the following:

  • [ ] Consistency: Does your robots.txt align with your actual server responses? (e.g., Are you allowing a path in robots.txt that your WAF blocks with a 403?)
  • [ ] Security: Are all private, administrative, and user-specific routes protected by actual server-side authentication (Layer 3), rather than just Disallow rules?
  • [ ] Visibility: Are you accidentally blocking AI search crawlers (e.g., OAI-SearchBot) at the WAF layer (Layer 4) due to overly aggressive anti-bot rules?
  • [ ] Data Protection: Are you using X-Robots-Tag (Layer 2) to prevent the indexing of non-HTML assets like PDFs and images?
  • [ ] Monitoring: Are you regularly reviewing your WAF and server logs to identify discrepancies between declared intent and actual bot behavior?

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.

Related policies