Bot directory / AI Assistant

ClaudeBot — Anthropic crawler

A technical reference for ClaudeBot and assistant-facing access policy.

AI Summary: ClaudeBot is catalogued as an assistant crawler. An explicit policy block can be evaluated, but it is not a contractual guarantee of downstream behavior.

Role and policy boundary

Treat this User-Agent as a distinct policy identity. The correct decision depends on the bot purpose, the paths it requests, and the response controls applied by the origin, CDN, and application layer. A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting.

ClaudeBot is grouped with assistant-oriented crawlers because an assistant may retrieve a page in response to a user request. That use case is different from training and from broad search indexing. Teams should decide whether public documentation is available to assistants while private, draft, or account-only pages remain out of scope.

An allow policy for public documentation can look like this:

configuration / code
User-agent: ClaudeBot
Allow: /docs/
Allow: /guides/
Disallow: /

The more specific public paths are intended to be reviewed together with the fallback Disallow. Keep the private boundary explicit:

configuration / code
User-agent: ClaudeBot
Disallow: /account/
Disallow: /billing/
Disallow: /internal/

Do not place credentials or personal data in a publicly reachable path and assume robots.txt will protect it. Robots rules are a preference signal, not an access-control system.

Layered verification

Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.

Suppose robots.txt allows /docs/, but the documentation page contains:

configuration / code
<meta name="robots" content="noindex">

or the origin returns:

configuration / code
X-Robots-Tag: noindex

The diagnostic should explain the contradiction: crawler-level access is open for the path, but page-level discovery is restricted. The fix depends on intent. If the documentation should be available to assistants, remove the unintended noindex directive; if the page is intentionally private, keep the restrictive layer and change the robots policy copy to match.

WAF and Nginx remediation examples

A WAF rule can first observe assistant traffic and preserve the request for analysis:

configuration / code
{
  "description": "Observe ClaudeBot assistant requests",
  "expression": "lower(http.user_agent) contains \"claudebot\"",
  "action": "log",
  "note": "Review public documentation paths before blocking."
}

When the decision is to block private paths at the edge, a Nginx location rule can be scoped narrowly:

configuration / code
location ~ ^/(account|billing|internal)/ {
    if ($http_user_agent ~* "ClaudeBot") { return 403; }
    try_files $uri $uri/ =404;
}

Nginx if directives require careful testing in production configurations. Prefer a tested map or existing access-control module when your platform provides one. The generated AI Bot Check snippet is advisory and should be reviewed by the owner of the edge stack.

Review checklist

Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.

Test the actual URL with the live fetch workflow, inspect the status code and X-Robots-Tag, and compare the fetched robots.txt with the page source. Record a limitation if the target blocks automated fetches or if the page is rendered entirely client-side. A partial result is useful evidence, but it is not proof of a pass.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.