Claude-Web: Robots.txt & Crawl Policy Reference
Technical reference for the registry-listed Claude-Web token and Anthropic's current ClaudeBot, Claude-User, and Claude-SearchBot policy distinctions.
AI Summary:
Claude-Web/1.0is a legacy or registry label that should not be assumed to represent a separate current Anthropic crawler. Anthropic’s current public guidance distinguishesClaudeBotfor model-development data collection,Claude-Userfor user-directed retrieval, andClaude-SearchBotfor search-result quality. The exactClaude-Webtoken is not listed in that current three-bot guidance, so verify the observed request and apply the policy for the documented bot that actually appears in logs.
Role and policy boundary
Anthropic’s current crawler guidance separates three purposes that are easy to conflate. ClaudeBot collects web content that could contribute to model development. Claude-User retrieves websites when an individual asks Claude a question. Claude-SearchBot navigates the web to improve search-result quality. These purposes lead to different visibility decisions: blocking model-development collection is not the same as blocking user-directed retrieval, and blocking search indexing is not the same as blocking every Anthropic request.
The registry contains a Claude-Web/1.0 label, but the official Anthropic guidance reviewed for this profile names the three bots above rather than a distinct Claude-Web category. Treat the legacy label as an observation that needs confirmation. Do not use it to make a broad claim about training, search, or user-triggered access. First identify the exact token in your edge logs; then apply the narrowest rule that matches the documented Anthropic purpose.
A robots rule is a published crawl preference; it does not authenticate a caller or protect private data. Anthropic states that its documented bots honor standard robots.txt directives and supports the non-standard Crawl-delay extension. That statement should be applied to the documented bot token, not automatically generalized to an unconfirmed Claude-Web token.
If your logs genuinely contain the legacy token and your policy is to block that declared identifier, use:
User-agent: Claude-Web
Disallow: /
If your objective is to block model-development collection, target the documented bot instead:
User-agent: ClaudeBot
Disallow: /
If you want to keep user-directed retrieval or search visibility separate, create separate groups for Claude-User and Claude-SearchBot according to the content owner’s decision. Do not replace these groups with a wildcard unless the same rule is intentionally meant for all crawlers.
Layered verification
Start with the exact User-Agent, path, method, status, redirects, request rate, and source address. Compare the observed source with Anthropic’s current machine-readable crawler IP list at https://claude.com/crawling/bots.json. Anthropic’s guidance says that addresses in this list indicate that the crawler is coming from Anthropic, which makes the list a useful corroborating signal. It is not a substitute for the request log, because addresses and bot contracts can change.
Next, evaluate the top-level /robots.txt from the canonical host. Check whether the file is publicly reachable, whether the dedicated group matches the actual token, and whether Crawl-delay is present where a slower request rate is required. Test the final redirected path rather than assuming that the initial URL has the same policy.
The Policy Engine should preserve the purpose distinction. A ClaudeBot deny result is a model-development opt-out; a Claude-User deny result can reduce visibility for user-directed questions; and a Claude-SearchBot deny result can reduce search indexing. A Claude-Web match without current vendor documentation should be reported as an evidence-limited legacy label, not silently mapped to all three behaviors.
Page-level directives are a separate signal:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives communicate indexing preferences but do not replace authentication, authorization, WAF policy, or rate limiting. Use them with a clear understanding of which consuming service honors them.
WAF and Nginx remediation examples
If your policy is to block only the legacy token observed in logs, a narrow WAF match is appropriate. It is not an identity check and should not be used as the sole protection for confidential content:
{
"description": "Block declared legacy Claude-Web token",
"expression": "lower(http.user_agent) contains \"claude-web\"",
"action": "block"
}
For a path-specific Nginx control:
map $http_user_agent $deny_claude_web {
default 0;
~*claude-web 1;
}
server {
location ~ ^/(internal|account|customer-data)/ {
if ($deny_claude_web) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not block all Anthropic traffic when the business decision concerns only one purpose. If you need a model-development opt-out, use ClaudeBot; if you need to reduce user-directed retrieval, use Claude-User; and if you need to control search indexing, use Claude-SearchBot. Keep private data protected by application authorization even when a robots group is permissive.
Review checklist
Request /robots.txt from every canonical subdomain and verify the exact group for the token observed in logs. Test one public documentation URL, one private URL, and one redirecting URL. Record the User-Agent, source address, status, final URL, headers, response size, and request rate. If using Anthropic’s crawler IP list as corroboration, record its retrieval date because the machine-readable list can change.
Finally, classify the outcome by purpose rather than by the word “Claude.” Confirm that the request is actually Claude-Web or update the policy to the documented current token. Re-test after CDN, WAF, origin, or robots changes. Do not claim that Claude-Web is a current Anthropic bot, that it trains Claude, or that it respects robots.txt unless fresh primary evidence supports that exact claim.
References
- Anthropic crawler policy — current public distinction between ClaudeBot, Claude-User, and Claude-SearchBot.
- Anthropic crawler IP list — machine-readable source prefixes and addresses used as a corroborating identity signal.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.