lb-spider: Robots.txt & Crawl Policy Reference
Technical reference for the historical lb-spider registry label, with explicit limits around its undocumented User-Agent, operator, and current crawl policy.
AI Summary:
lb-spideris a historical registry label for an LB spider used for content discovery, but no exact User-Agent or operator-owned documentation is available in the current inventory. A headless Bing search found no relevant operator source. Treat any matching traffic as unidentified until logs and an authoritative source establish the client and its policy.
Role and policy boundary
The inventory describes lb-spider as an LB spider web crawler for content discovery, but it provides no User-Agent value and links only to a community-maintained crawler registry. A headless Bing search for lb-spider crawler official returned no operator-owned domain or relevant documentation in the extracted results; visible results were unrelated pages. This is a research limitation, not evidence that the client never existed or is inactive.
The content-discovery description is therefore a registry hypothesis only. A request may come from an old deployment, an independent crawler, a partner tool, a fork, or a spoofed header. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private content from the label.
Because no exact token is documented, do not publish a made-up User-agent: lb-spider group as if it will reliably match the client. If logs later reveal a complete and independently attributable header, use that exact value:
User-agent: CONFIRMED-OBSERVED-TOKEN
Disallow: /
For selective access after confirmation:
User-agent: CONFIRMED-OBSERVED-TOKEN
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
These are defensive examples, not recovered LB spider instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No exact lb-spider token or operator source-verification procedure was available in this review. Do not reuse rules for another spider, search engine, or hosting provider.
Classify request behavior without turning it into attribution. Public HTML, canonical metadata, feeds, and sitemaps may be consistent with discovery. Private endpoint access, high concurrency, repeated retries, unexpected downloads, or traffic that ignores your policy may indicate spoofing, abuse, a fork, or another client. These observations establish operational impact, not operator identity or downstream use.
Evaluate /robots.txt only after the actual header is known. Confirm that it is served by the correct host, returns a successful status and text content type, and contains the exact group you intend to publish. Test the full observed token and any global group separately. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate an undocumented spider or secure private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
When the identity is unknown, use a report-only candidate rule and do not block broad spider or LB substrings. Adapt the expression to your WAF provider:
{
"description": "Log unidentified LB spider candidates",
"expression": "lower(http.user_agent) contains \"lb-spider\"",
"action": "log"
}
Once logs and an authoritative source establish a complete token, replace the candidate with a narrow route-scoped control:
map $http_user_agent $block_confirmed_lb_spider {
default 0;
# Add only a complete, independently verified token here.
# ~*Confirmed-LB-Spider-Token 1;
}
server {
location ~ ^/(admin|account|private|internal|api)/ {
if ($block_confirmed_lb_spider) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not invent an IP allowlist, reverse-DNS suffix, or permanent deny rule from the registry name. A User-Agent is easy to spoof and the candidate expression may catch unrelated clients. Test public HTML, feeds, sitemaps, media, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for broad lb and spider candidates only to locate a complete header, then preserve the exact value. Record source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, and redirect chain. Check whether a current operator source identifies the client before assigning purpose or a robots rule.
Re-check the community registry and run a fresh search when a credible operator domain becomes discoverable. Keep this profile at legacy-label and the User-Agent as Not publicly documented until a first-party source publishes a current token, purpose, source-verification method, or robots guidance. Treat the unsuccessful headless search as a source-discovery limitation, not proof of inactivity.
Decide whether your objective is to preserve discovery, limit extraction, protect private material, or reduce crawl load. Publish exact robots rules only after the token is known, and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent logs and response behavior.
References
- Crawler User Agents community registry — registry-linked community source; it provides no operator-owned lb-spider documentation in the inventory.
- Headless Bing search for lb-spider — search returned no relevant operator-owned source in the extracted results on 2026-08-25.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.