crawl4ai: Robots.txt & Crawl Policy Reference
Technical reference for Crawl4AI and the registry-listed crawl4ai-adapter/1.0 token. Learn why software-configurable User-Agents require request-level verification.
AI Summary: Crawl4AI is an open-source, LLM-friendly web crawler and scraper for RAG, agents, and data pipelines. Its official documentation supports custom browser User-Agents and a crawl option to check and respect
robots.txt, so it is not one fixed vendor-operated crawler. The registry-listedcrawl4ai-adapter/1.0token should be treated as an observed identifier; site owners must audit the actual request rather than assuming that every Crawl4AI user shares the same policy behavior.
Role and policy boundary
Crawl4AI is software that developers run and configure. The project describes itself as an open-source web crawler and scraper that converts web content into clean, LLM-ready Markdown for RAG, agents, and data pipelines. Its documentation includes controls for browser type, custom User-Agent, headers, cookies, sessions, proxies, and crawling behavior. That makes it fundamentally different from a single managed crawler with one stable operator identity.
A Crawl4AI deployment may be used to extract documentation for a private RAG system, collect public data for research, test a website, or build a broader ingestion pipeline. The project’s purpose does not reveal the intent of the person running it. The same library can be configured with a browser-like User-Agent, a project-specific token, or a custom identifier. The registry token crawl4ai-adapter/1.0 is therefore a signal to investigate, not proof that a request is official Crawl4AI traffic.
Crawl4AI’s crawler configuration documents a setting to check and respect robots.txt before crawling. That is a capability available to the operator; it is not a guarantee that every deployment enables it or that a caller has no other access path. A robots rule communicates the site owner’s preference but does not authenticate the user running the software, protect private URLs, or reveal how retrieved content will be used.
If you have observed the registry token and want to block only that declared identifier, use a narrow group:
User-agent: crawl4ai-adapter
Disallow: /
If you want to permit public documentation but exclude private and high-cost routes, use:
User-agent: crawl4ai-adapter
Allow: /docs/
Allow: /public-data/
Disallow: /staging/
Disallow: /internal/
Disallow: /account/
Disallow: /checkout/
Do not block every browser-like request or every request from a shared hosting provider based only on the library name. Protect confidential data with application authorization and use robots, WAF, and rate limits for their appropriate purposes.
Layered verification
Start with the complete request record: User-Agent, source address, method, path, timestamp, status, redirect chain, response size, and request rate. Compare the observed identifier with the registry value, but remember that Crawl4AI supports custom User-Agents. A request that says crawl4ai-adapter/1.0 may be produced by a user running the library, a wrapper around the library, or unrelated software imitating the token.
Next, evaluate the canonical top-level /robots.txt and the exact path scope. Crawl4AI’s documentation indicates that an operator can configure the crawler to check and respect robots.txt, but the site owner cannot infer from a User-Agent alone whether that option is enabled. Test with your own controlled endpoint if you operate the crawler, and rely on access controls rather than an honor-system robots rule for sensitive data.
Inspect HTML metadata and response headers as independent signals. A noindex response can express a discoverability preference, but it does not prevent a request. A WAF deny can stop a connection, but it does not remove the URL from a public response already served to another client. The Policy Engine should preserve these distinctions and mark the bot identity as configurable rather than vendor-authenticated.
For page-level indexing preferences, review:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives do not replace authentication, authorization, WAF policy, or rate limiting. If the business concern is AI training or RAG ingestion, document the intended policy separately from ordinary search indexing.
WAF and Nginx remediation examples
If your logs show the registry token and you decide to block it, a narrow WAF match can reduce repeat requests. Because the software is configurable, record this as a self-declared-token match and not as proof of the caller’s identity:
{
"description": "Block declared Crawl4AI adapter token",
"expression": "lower(http.user_agent) contains \"crawl4ai-adapter\"",
"action": "block"
}
For a selective path rule in Nginx:
map $http_user_agent $deny_crawl4ai {
default 0;
~*crawl4ai-adapter 1;
}
server {
location ~ ^/(internal|customer-data|account)/ {
if ($deny_crawl4ai) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not use this User-Agent match as the only protection for private content. If you control a Crawl4AI deployment, configure its robots handling explicitly, set a truthful identifying User-Agent, apply rate limits, and honor the site owner’s policy. If you operate the site, use authentication and origin authorization when the data must not be retrieved.
Review checklist
Request /robots.txt from the canonical host and verify its status, content type, final URL, and exact rule for the observed token. Test one public documentation URL, one disallowed internal URL, and one protected account URL. Record the complete User-Agent, source address, HTTP status, redirects, response headers, response size, and request rate.
If you operate the crawler, confirm whether the robots-check option is enabled and whether the configured User-Agent accurately identifies your application. If you are defending the site, do not assume that all Crawl4AI traffic uses one token. Re-test after CDN, WAF, origin, or crawler-configuration changes, and keep the evidence boundary visible in the report.
References
- Crawl4AI Browser, Crawler & LLM Config — official configuration reference, including custom User-Agent and robots.txt handling.
- Crawl4AI official repository — open-source project scope and crawler/scraper capabilities.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.