grub.org: Robots.txt & Crawl Policy Reference
Technical reference for the historical grub-client crawler label, with explicit limits around current activity, operator documentation, and request verification.
AI Summary:
grub-client-0.3.0is a historical registry User-Agent associated with the Grub search crawler. Direct requests togrub.organdwww.grub.orgfailed withnet::ERR_CONNECTION_CLOSEDduring review, so no current operator policy, robots behavior, IP range, or verification method was established. Treat matching traffic as an unverified historical observation.
Role and policy boundary
The inventory associates grub.org with a Grub search-engine crawler and records this historical User-Agent:
Mozilla/4.0 (compatible; grub-client-0.3.0; Crawl your own stuff with http://grub.org)
The linked source is a community-maintained crawler registry rather than a current Grub operator policy. Direct headless-browser requests to both https://grub.org/ and https://www.grub.org/ closed the connection, and an HTTP attempt was interrupted after the same connectivity problem. These failures do not prove that Grub never existed or that every historical client is inactive; they mean that a current first-party source was not available in this review.
Treat search indexing as the registry’s historical role hypothesis only. A matching header may come from a legacy deployment, an independent client, a fork, or a spoofed request. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private material from the token.
If your logs confirm this exact token and your site policy is to exclude it, a narrow defensive rule could be:
User-agent: grub-client-0.3.0
Disallow: /
For selective access to approved public pages:
User-agent: grub-client-0.3.0
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
These are site-owner examples, not recovered Grub instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No current Grub source reviewed for this profile published a network range, DNS verification procedure, or rate policy. Do not treat the header alone as authentication.
Compare observed behavior with a search-spider hypothesis without turning it into attribution. Requests for public HTML, canonical metadata, feeds, sitemaps, and ordinary assets may be consistent with indexing. High concurrency, private endpoint access, repeated retries, unexpected downloads, or traffic that ignores your site policy may indicate spoofing, abuse, a fork, or a different client. These observations establish impact, not operator identity or downstream use.
Evaluate /robots.txt independently if your site chooses to publish a rule. Confirm that the response is served from the intended host, returns a successful status and text content type, and contains the exact group you intend to apply. Test the full token and any global group separately. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate a historical Grub client or secure private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If logs show a repeatable unwanted token, start in report-only mode and preserve representative requests. Adapt the expression to your WAF provider; it identifies a historical header pattern only:
{
"description": "Review historical grub-client traffic",
"expression": "lower(http.user_agent) contains \"grub-client-0.3.0\"",
"action": "log"
}
After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:
map $http_user_agent $review_grub_client {
default 0;
~*grub-client-0\.3\.0 1;
}
server {
location ~ ^/(admin|account|private|internal|api)/ {
if ($review_grub_client) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not invent an IP allowlist, reverse-DNS suffix, or current operator exception. A User-Agent is easy to spoof and a broad historical match can catch an authorized test or unrelated client. Test public HTML, feeds, sitemaps, media, uploads, account flows, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for the complete registered header and preserve representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the source can be independently verified and whether traffic matches a public-indexing pattern. Do not classify a request as authentic from its token or the community registry entry.
Re-check grub.org, www.grub.org, and the registry for a current operator source when the service becomes reachable. Keep this profile at legacy-label until a credible source publishes a current policy, canonical User-Agent, source-verification method, or robots guidance. Treat the connection-closed result as a source-access limitation, not proof of inactivity.
Decide whether your objective is to preserve search visibility, limit extraction, protect private material, or reduce crawl load. Publish exact robots rules and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent access logs and response behavior.
References
- Grub registry source — community-maintained crawler User-Agent list linked by the inventory; it is not a current Grub operator policy.
- grub.org — direct headless-browser request failed with
net::ERR_CONNECTION_CLOSEDon 2026-08-25. - www.grub.org — direct headless-browser request failed with
net::ERR_CONNECTION_CLOSEDon 2026-08-25. - Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.