Bot directory / scraper

curl: Robots.txt & Crawl Policy Reference

Technical reference for curl, a command-line HTTP client commonly used in scripts that does not natively respect robots.txt.

AI Summary: curl is a standard, open-source command-line HTTP client used globally for automation, API testing, and scraping. It does not natively parse or respect robots.txt. Because anyone can use curl from any IP address, there is no official operator IP range. Use robust network validation, rate limiting, and behavioral analysis to manage curl traffic rather than relying on robots.txt.

Role and policy boundary

curl is a fundamental command-line tool and library for transferring data with URLs. The registry records a typical User-Agent as curl/8.7.1 (the version number varies). The official documentation at https://curl.se/ details its extensive capabilities.

Because curl is a generic tool, it does not natively fetch, parse, or respect robots.txt. Any script using curl must implement its own robots.txt compliance logic, which is rarely done. Furthermore, since anyone can run curl, there is no specific operator, IP range, or reverse-DNS procedure to verify. This profile is marked documented-limit because while the tool is well-documented, the source of any specific request cannot be authenticated via the User-Agent.

Adding curl to your robots.txt will generally be ignored by the tool itself:

configuration / code
User-agent: curl
Disallow: /

These are explanatory site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls. To effectively block unwanted curl traffic, you must use WAF rules, rate limiting, or application-level blocks.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. There is no official curl IP range or verification workflow. The User-Agent is useful for triage but is not authentication.

Do not treat the word curl in a header as proof of benign intent. Check source IP and DNS evidence independently, and require behavioral consistency before creating a trust exception. Many automated attacks, scrapers, and vulnerability scanners use the default curl User-Agent.

Compare observed behavior with a generic script hypothesis without turning it into attribution. Public HTML, metadata, feeds, sitemaps, and ordinary assets may be requested by many tools. Private endpoints, licensed content, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not curl identity or downstream use.

Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish. An absent IP list is not evidence of robots non-compliance, but it prevents proactive allowlisting.

Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate an undocumented crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes search indexing, AI input, reference use, and model training, document each purpose separately rather than inferring permission from a registry label.

WAF and Nginx remediation examples

Because the identity cannot be cryptographically verified via published IP lists, use report-only logging and a narrow observation. Avoid broad curl rules that can affect legitimate API clients or internal health checks:

configuration / code
{
  "description": "Observe curl candidates",
  "expression": "lower(http.user_agent) contains \"curl/\"",
  "action": "log"
}

After independent verification and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule. If you decide to block default curl User-Agents from web routes:

configuration / code
map $http_user_agent $block_curl_private {
    default 0;
    ~*^curl/ 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal)/ {
        if ($block_curl_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof or change (e.g., curl -A "Mozilla/5.0..."). Do not invent an IP allowlist, ASN exception, reverse-DNS suffix, rate, training policy, or permanent trust exception based solely on this header. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact curl substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. The operator does not provide a verification method.

Re-check the crawler registry, the curl project site, robots.txt, and operator contact path. Keep this profile at documented-limit as it is a generic tool without a specific operator IP space.

Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.

Review search indexing, AI input, reference use, and model-training decisions separately. The documentation does not establish downstream permission or use. Do not claim successful blocking or verification from configuration alone; validate later logs.

Record the date and reason for the documented-limit classification so future evidence can be compared. If a new deployment identifies itself differently, create a separate evidence trail.

References

  1. curl Project Official Site — official documentation for the curl command-line tool.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.