DeepSeekBot: Robots.txt & Crawl Policy Reference
Technical reference for the registry-listed DeepSeekBot token. Compare DeepSeek's official product and API documentation with observed request evidence before applying crawler controls.
AI Summary:
DeepSeekBot/1.0is a registry-listed User-Agent associated with DeepSeek, but the official DeepSeek company site and API documentation reviewed for this profile do not publish a definitive public crawler specification or robots.txt compliance statement for this exact token. Treat it as an unverified observation, verify the request in logs, and do not claim model-training use or operator identity from the User-Agent alone.
Role and policy boundary
DeepSeek’s official site identifies the company as an AI research organization and links to its web product, API platform, models, and developer documentation. The official API documentation covers API endpoints, model access, and agent integrations. Those sources establish the vendor and its AI product context, but they do not establish that a public crawler called DeepSeekBot is currently operating or explain what a request carrying this token would do with retrieved content.
The registry classifies DeepSeekBot as an AI model-training and data-collection crawler. That classification should remain a registry hypothesis until a primary DeepSeek source or reliable network evidence confirms it. A User-Agent is self-declared and can be copied by unrelated clients. A browser request made by a DeepSeek user, an API integration, or an independent data collector may not use the same token at all.
For a site owner, the right first step is to observe rather than assume. Record the exact User-Agent, source address, path, method, status, redirect chain, response size, and request rate. If the token is genuinely present and the organization does not want the identified request pattern to access public content, a dedicated robots group can express the preference:
User-agent: DeepSeekBot
Disallow: /
For a selective policy that allows public documentation while excluding sensitive paths, use:
User-agent: DeepSeekBot
Allow: /docs/
Disallow: /staging/
Disallow: /internal/
Disallow: /account/
Disallow: /customer-data/
A robots rule is a crawl preference, not authentication. It does not protect a private URL that is already public, prevent an authenticated API workflow from receiving data, or prove the downstream use of a request.
Layered verification
Begin with the raw request evidence and compare it with the current DeepSeek product and API pages. The reviewed primary sources do not publish the exact DeepSeekBot/1.0 token or a crawler IP list, so the registry value should be marked as unverified. Do not reverse-DNS a source address into a vendor attribution without corroborating it through current operator documentation.
Then evaluate the canonical top-level /robots.txt. Confirm that the file is publicly reachable, returns the expected status and content type, and contains a group matching the exact observed token. Test the actual requested path, not just the site root, and record any redirect or CDN rewrite that changes the effective URL.
The Policy Engine should preserve the evidence boundary. A match to DeepSeekBot can explain which site rule was selected, while the absence of a primary crawler specification remains visible. Do not turn DeepSeek’s documented AI models or API capabilities into a claim that the company is collecting a particular page for training.
Inspect page-level metadata and response headers separately:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals communicate indexing preferences but do not provide access control. Use authentication, application authorization, WAF rules, and rate limiting for private data and resource protection.
WAF and Nginx remediation examples
If the token is observed and your policy is to block that self-declared identifier, a narrow WAF expression can reduce repeat requests. It is a network-layer match, not proof that the request is operated by DeepSeek:
{
"description": "Block declared DeepSeekBot token",
"expression": "lower(http.user_agent) contains \"deepseekbot\"",
"action": "block"
}
A path-specific Nginx example is:
map $http_user_agent $deny_deepseekbot {
default 0;
~*deepseekbot 1;
}
server {
location ~ ^/(internal|customer-data|account)/ {
if ($deny_deepseekbot) { return 403; }
try_files $uri $uri/ =404;
}
}
Test the rule in staging and do not use a User-Agent match as the sole control for confidential content. Do not block all DeepSeek API traffic or all traffic from a hosting network based on this unverified token. If DeepSeek publishes an official crawler policy, update the exact token, purpose, and source date together.
Review checklist
Request /robots.txt from each canonical subdomain and verify the status, content type, final URL, and exact rule applied to the observed token. Test one public documentation URL, one disallowed internal URL, and one redirecting URL. Record the complete User-Agent, source address, HTTP status, redirects, response headers, response size, and request rate.
Re-check DeepSeek’s official site and API documentation for a current crawler notice. If the token is not observed, leave the profile marked as unverified rather than treating the registry entry as proof. Re-test after CDN, WAF, origin, or robots changes and preserve the evidence needed to distinguish DeepSeek product traffic from an unrelated client.
References
- DeepSeek official site — company, web product, API platform, and model context.
- DeepSeek API documentation — official API and agent-integration documentation; no fixed public crawler specification is published on the reviewed page.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.