Cloudflare-AutoRAG: Robots.txt & Crawl Policy Reference
Technical reference for the registry-listed Cloudflare-AutoRAG token and Cloudflare's documented AI Search and AutoRAG ingestion workflows.
AI Summary:
Cloudflare-AutoRAGis a registry-listed User-Agent associated with Cloudflare’s AutoRAG terminology. Cloudflare’s official documentation describes AI Search as a managed search primitive and AutoRAG as an ingestion, indexing, retrieval, and generation workflow. The reviewed primary sources do not independently confirm a public crawler contract for this exact token, so treat it as an observed identifier and do not assume that every AutoRAG data source is collected through a public web crawl.
Role and policy boundary
Cloudflare documents AutoRAG as a managed Retrieval-Augmented Generation workflow with two broad stages: asynchronous indexing and synchronous querying. The documented pipeline can ingest files from a data source, convert content to Markdown, chunk it, create embeddings, store vectors, and retrieve relevant content for a generated response. Cloudflare AI Search is described as a search primitive for applications and agents that can index connected data and query it with natural language.
That product description does not establish that all AutoRAG instances crawl the public web. A deployment may ingest content from R2, an application pipeline, an API, or a website-parsing workflow using Cloudflare services. The registry-listed Cloudflare-AutoRAG token therefore needs to be separated from the product name. A website owner should not infer the data source, retention, or downstream use of a request solely from this User-Agent.
If the token appears in your edge logs and you do not want that public path collected, a dedicated robots group can express the preference:
User-agent: Cloudflare-AutoRAG
Disallow: /
If you want to expose a public documentation area but protect internal and transactional routes, use a path-specific policy:
User-agent: Cloudflare-AutoRAG
Allow: /docs/
Allow: /public-knowledge/
Disallow: /staging/
Disallow: /internal/
Disallow: /account/
Disallow: /checkout/
A robots rule is a crawl preference, not authentication. It does not prevent a Cloudflare customer from ingesting content through an authenticated API, a connector, or an application-side workflow, and it does not protect private content that is already publicly accessible. Use authorization at the origin for confidential data.
Layered verification
Start by recording the complete request evidence: User-Agent, source address, path, method, status, redirects, response size, and request rate. The registry value includes https://developers.cloudflare.com/autorag and an email-like contact string, but the current Cloudflare AI Search and AutoRAG sources reviewed here do not independently publish that exact crawler token. Treat the header as a self-declared observation rather than vendor authentication.
Next, test the canonical top-level /robots.txt, including the exact hostname that received the request. Verify that the file returns the intended content without a login page, challenge, or unexpected redirect. Compare the path-specific robots result with HTML metadata and HTTP response headers. If the request belongs to a configured Cloudflare workflow, confirm the actual data source and ingestion method in the account configuration rather than inferring it from the User-Agent.
The Policy Engine evaluates these signals independently. A match to Cloudflare-AutoRAG explains which robots group was selected, while the lack of a public crawler specification remains an evidence limitation. A permissive robots rule does not prove that a request came from Cloudflare, and a page-level noindex directive does not prevent an authenticated application from retrieving the page.
For pages that should not be indexed or offered as public knowledge, review separate page-level controls:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives communicate discoverability and indexing preferences. They do not replace authentication, WAF controls, rate limiting, or removal of confidential material from a public response.
WAF and Nginx remediation examples
If you choose to block the self-declared token, use a narrow WAF rule and record that it is a network-layer match rather than an identity proof:
{
"description": "Block declared Cloudflare-AutoRAG token",
"expression": "lower(http.user_agent) contains \"cloudflare-autorag\"",
"action": "block"
}
For a path-specific Nginx rule that protects internal knowledge routes:
map $http_user_agent $deny_cloudflare_autorag {
default 0;
~*cloudflare-autorag 1;
}
server {
location ~ ^/(internal|customer-data|account)/ {
if ($deny_cloudflare_autorag) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not use this User-Agent match as the only safeguard for confidential content, and do not block all Cloudflare traffic when the policy concerns only a registry-listed token. If the desired outcome is to keep public documentation searchable while excluding private sources, combine path-specific robots guidance with application authorization and documented ingestion settings.
Review checklist
Request /robots.txt from the canonical host and verify the exact Cloudflare-AutoRAG group, status, content type, final URL, and path scope. Test one allowed documentation URL, one disallowed internal URL, and one response containing page-level metadata. Record the User-Agent, source address, HTTP status, redirects, response size, and request rate.
Then confirm whether the observed request corresponds to a public crawler, a configured website parser, or another Cloudflare product workflow. Review the current Cloudflare AI Search and AutoRAG documentation after product changes. Keep the evidence classification visible: the product and ingestion behavior are documented, but the exact public crawler token remains unconfirmed in the sources reviewed here. Re-test after CDN, WAF, origin, or content-policy changes.
References
- Cloudflare AI Search documentation — managed search, indexing, retrieval, and website data-source concepts.
- Introducing AutoRAG on Cloudflare — official description of ingestion, indexing, embeddings, vector storage, and querying.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.