Scrubby: Robots.txt & Crawl Policy Reference
Technical reference for the Scrubby registry label, distinguishing the active Scrub the Web SEO directory from an undocumented and unverified crawler identity.
AI Summary: The registry describes
Scrubbyas a Scrub the Web business-directory crawler and recordsScrubby/2.2 (http://www.scrubtheweb.com/). The linked domain hosts a live SEO business directory, but the current first-party pages do not document a crawler, User-Agent, robots policy, IP range, or verification method. Treat the token as an unverified historical label rather than an authenticated crawler identity.
Role and policy boundary
The registry describes this entry as a Scrub the Web business directory crawler and records this User-Agent:
Scrubby/2.2 (http://www.scrubtheweb.com/)
The registry-linked domain currently resolves to an active SEO-friendly business directory offering paid submissions for backlinks. However, the reviewed first-party pages—including the homepage and About page—do not publish a crawler specification, explain its purpose, or document the Scrubby User-Agent.
This profile therefore uses legacy-label and does not present the registry description as a current operator policy. A matching request could be a legacy deployment, an undocumented directory verification tool, a test client, or a spoofed header. Do not infer that the directory currently operates a broad crawler, that it obeys robots.txt, or that the token establishes permission for AI input, model training, or data retention.
Because no current policy was verified, do not publish an assumed Scrubby robots group as if it came from the operator. If logs later establish a current, attributable token and you decide to exclude it, use the exact observed value in a deliberate site-owner rule:
User-agent: CONFIRMED-OBSERVED-SCRUBBY
Disallow: /
For selective access after confirmation:
User-agent: CONFIRMED-OBSERVED-SCRUBBY
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/
These are explanatory site-owner controls, not recovered Scrub the Web instructions. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. No current source in this review published a Scrubby IP range, reverse-DNS procedure, Crawl-delay, rate guidance, removal address, or verification workflow.
Do not treat the scrubtheweb.com URL in a header as proof of ownership. A client can copy any URL. Check source IP and DNS evidence independently, and require a documented current verification path before creating a trust exception. An IP that happens to belong to the directory’s hosting provider is not alone proof that a request is an authorized crawler.
Compare observed traffic with a directory-indexing hypothesis without turning it into attribution. Public HTML, metadata, feeds, sitemaps, and static assets may be requested by many tools. Private endpoints, licensed content, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not directory identity or downstream use.
Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish. An active directory website does not establish that its namesake crawler obeys robots rules. If no exact group exists, a global rule may affect unrelated clients and should be adopted only as an explicit site-wide decision.
Page-level directives can express indexing preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate an undocumented crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes directory indexing, AI input, reference use, and model training, document each purpose separately rather than inferring permission from a registry label.
WAF and Nginx remediation examples
When the identity is unverified, use report-only logging and a narrow observation. Avoid broad scrub, directory, or browser rules that can affect unrelated clients:
{
"description": "Observe unverified Scrubby candidates",
"expression": "lower(http.user_agent) contains \"scrubby\"",
"action": "log"
}
After independent verification and a policy decision, scope enforcement to sensitive routes and preserve evidence supporting the rule:
map $http_user_agent $block_confirmed_scrubby_private {
default 0;
# Add only a complete, independently verified token here.
# ~*Scrubby/2\.2 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_confirmed_scrubby_private) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof and the registry token is undocumented. Do not invent an IP allowlist, directory ASN exception, reverse-DNS suffix, rate, training policy, or permanent trust rule. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.
Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.
Review checklist
Search logs for the exact Scrubby substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. Treat the URL in the header as a clue, not proof of ownership.
Re-check the Scrub the Web directory, the registry source, robots.txt, and any future operator contact path. In this review the corporate website was accessible but did not document the crawler. Keep the profile at legacy-label until a current source publishes an exact token, purpose, robots behavior, source-verification method, rate guidance, or opt-out process.
Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.
Review directory indexing, AI input, reference use, and model-training decisions separately. Neither the registry description nor the current directory page establishes downstream permission or use. Do not claim successful blocking or verification from configuration alone; validate later logs and obtain a current operator response when possible.
Record the date and reason for the legacy-label classification. If a future deployment identifies itself with a concrete token, create a new evidence trail and update policy only after confirming that source and network.
References
- Scrub the Web — current first-party directory context; reviewed with the headless browser on 2026-08-25, without crawler documentation.
- Crawler User Agents community registry — registry context for the Scrubby label; no current operator documentation was established in this review.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.