Apache Nutch: Robots.txt & Crawl Policy Reference
Technical reference for Apache Nutch deployments, distinguishing the configurable open-source crawler framework from one unverified registry User-Agent and site-specific policy.
AI Summary: Apache Nutch is an open-source, extensible web-crawler framework maintained by the Apache Software Foundation, not one centrally operated crawler network with a single universal policy. The registry records a browser-like
Friendly_Crawler/2.0 ... Nutch-1.20-SNAPSHOTtoken, but Apache’s official pages reviewed here do not establish that exact header as a default or current Apache-controlled identity. Verify the deployed configuration, source IP, DNS, robots behavior, and purpose before allowing or blocking traffic.
Role and policy boundary
Apache’s official homepage describes Nutch as a highly extensible and scalable web crawler for a wide range of configurable data-acquisition tasks. It highlights batch processing with Hadoop and plugins for parsing, indexing, HTML filtering, and scoring, with integrations such as Solr and Elasticsearch. The official documentation index links to project information, FAQs, Javadoc, security material, tutorials, and a community wiki.
Those facts describe a software framework. They do not mean every Nutch deployment is operated by Apache, uses the same crawl schedule, indexes the same content, or has the same downstream data practices. A site administrator, research group, search service, or another product can configure the framework and choose its own User-Agent, request rate, seed list, storage, and plugins.
The registry stores this long browser-like value:
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/605.1.16 (KHTML, like Gecko; compatible; Friendly_Crawler/2.0) Chrome/120.0.6099.217 Safari/605.1.15/Nutch-1.20-SNAPSHOT
Apache’s reviewed pages do not establish that this exact value is a default or current Apache-controlled User-Agent. Preserve it as a registry candidate and compare it with the complete observed header. Do not merge it with a framework installation that identifies itself differently, and do not attribute all requests containing Nutch to the Apache project without deployment evidence.
If you operate a Nutch instance, publish a stable, descriptive User-Agent and a contact URL. If you are a site owner deciding what to do with an observed token, first identify the deploying organization and purpose. A robots group can be used after the token is known:
User-agent: CONFIRMED-NUTCH-DEPLOYMENT
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/
Crawl-delay: 5
For a deliberate exclusion:
User-agent: CONFIRMED-NUTCH-DEPLOYMENT
Disallow: /
These are site-owner examples, not a universal Apache Nutch instruction. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. No universal Apache Nutch IP range, reverse-DNS suffix, Crawl-delay, or operator contact was established by the reviewed official pages because those properties belong to each deployment.
Inspect the observed User-Agent for the exact framework/version string, but treat it as a clue. Check forward-confirmed reverse DNS only when an authoritative deploying organization publishes a hostname policy, and require the forward lookup to return the observed address. A header alone never authenticates a crawler. An IP that appears in one deployment’s documentation must not be generalized to every Nutch installation.
Compare the traffic with the claimed task. Public HTML, metadata, feeds, sitemaps, and ordinary assets may be consistent with search discovery, extraction, testing, or another batch job. Private endpoints, licensed content, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not a particular operator or downstream use.
Evaluate /robots.txt independently. Confirm that it is served by the intended host, returns a successful text response, and contains the exact User-agent group you intend to publish. Framework support for robots parsing does not prove that an arbitrary deployment obeyed the file. Check actual request behavior against your published paths, and use application authorization for sensitive resources.
Page-level directives can express indexing preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate Nutch or restrict a configured client that ignores them. If your data policy distinguishes search indexing, AI input, reference use, and model training, document each purpose separately; Nutch’s framework role does not decide downstream use.
WAF and Nginx remediation examples
When the deployment is not identified, use report-only logging and a narrow observation for the registry token. Do not block every browser-compatible request or every string containing Nutch:
{
"description": "Observe unverified Nutch deployment candidates",
"expression": "lower(http.user_agent) contains \"nutch\" or lower(http.user_agent) contains \"friendly_crawler\"",
"action": "log"
}
After identifying the deploying party and deciding to protect sensitive paths, scope enforcement to the exact documented token or a stronger identity signal:
map $http_user_agent $block_confirmed_nutch_private {
default 0;
# Add only a complete, independently verified deployment token here.
# ~*Friendly_Crawler/2\.0.*Nutch-1\.20-SNAPSHOT 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_confirmed_nutch_private) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof, and Nutch installations can be customized. Do not invent an Apache-owned IP allowlist, reverse-DNS suffix, global rate, training policy, or permanent trust exception. Use per-client rate limits, concurrency ceilings, timeouts, caching, response-size controls, and anomaly detection at the edge or origin. Begin in report-only mode, review false positives, and switch to a restrictive action only after evidence supports it.
Test public pages, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.
Review checklist
Record the exact observed User-Agent, source IP, ASN, PTR and forward lookup, requested paths, methods, response sizes, statuses, timing, concurrency, rate, redirects, and deployment contact. Compare the header to the registry candidate, but do not call it Apache-controlled without evidence from the deploying organization.
Identify whether the traffic is a specific Nutch installation, an Apache project test, a private research crawler, or an unrelated spoof. Confirm the deployment’s purpose and policy independently. The official Apache pages document the framework, not a universal network, current IP list, robots group, or operator behavior.
Check your robots file for an exact deployment group, test precedence and path behavior, and measure origin load. Use meta directives and X-Robots-Tag for indexing preferences, but use authentication, authorization, signed URLs, and WAF controls for private or licensed resources.
Review search indexing, AI input, reference use, and model-training decisions separately. Do not infer downstream permission or use from the framework name, a search-related deployment, or a browser-like token.
Keep the profile at documented-limit: Apache documents Nutch as a framework, while the registry’s exact User-Agent and any individual operator policy remain unverified. Re-check the official project documentation and the deploying organization’s source when the version or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.
References
- Apache Nutch — official project homepage describing the framework’s scalable, extensible web-crawler role and integrations; reviewed with the headless browser on 2026-08-25.
- Apache Nutch documentation index — official links to About, FAQs, Javadoc, Security, Tutorials, and Wiki material; reviewed with the headless browser on 2026-08-25.
- Apache Nutch source repository — official project source link; a framework repository does not authenticate every deployment or User-Agent.
- Crawler User Agents community registry — registry context for the long browser-like token; it does not establish a universal Apache-controlled identity.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.