lyonl-crawler: Robots.txt & Crawl Policy Reference
Technical reference for the lyonl-crawler registry label, including Lyonl’s current first-party crawler policy and explicit limits around the unlisted token.
AI Summary: The registry records
lyonl-crawler/0.1 (+https://lyonl.com/crawler; crawler@lyonl.com), but Lyonl’s current first-party policy listsLyonlBotas the canonical web crawler with a different User-Agent and robots token. Lyonl documents public search discovery, robots/Crawl-delay, bounded Chromium rendering, and a support path, while its IP endpoint has no published ranges. Treat the registry token as unlisted until Lyonl confirms it.
Role and policy boundary
The inventory describes lyonl-crawler as the Lyonl open-web search crawler and records this User-Agent:
lyonl-crawler/0.1 (+https://lyonl.com/crawler; crawler@lyonl.com)
The registry URL redirected to Lyonl’s current first-party policy at https://lyonl.com/bot.html. That policy lists the canonical identity as:
Mozilla/5.0 (compatible; LyonlBot/1.0; +https://lyonl.com/bot.html)
with the robots token lyonlbot. It also lists separate LyonlBot Image and LyonlBot News identities. The registry token lyonl-crawler/0.1 does not appear in the official identity table. Lyonl says that an unlisted Lyonl User-Agent should be reported with the IP address, timestamp, requested URL, and full header through its feedback path.
The current Lyonl policy describes the canonical crawler as visiting public URLs to build and maintain the Lyonl Search index. It may discover documents through links, sitemaps, redirects, and previously known URLs; it uses ETag and Last-Modified validators where available, and it does not intentionally bypass authentication, paywalls, CAPTCHAs, or access controls. It also says the crawler does not submit forms, register accounts, or perform transactions. These statements apply to the listed Lyonl service, not automatically to an unlisted token or a spoofed request.
Do not merge lyonl-crawler silently with LyonlBot. If Lyonl confirms that the registry token is a valid identity, use an exact robots rule:
User-agent: lyonl-crawler
Disallow: /
For selective access after confirmation:
User-agent: lyonl-crawler
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /private/
Disallow: /account/
Disallow: /api/
Crawl-delay: 10
These are site-owner examples, not recovered instructions for the registry token. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Lyonl’s policy warns that User-Agent strings can be spoofed and says it is still working on a public verification method. Until crawler IP ranges are published, Lyonl asks site owners to report the IP, UTC timestamp, URL, full header, and relevant access-log lines through its crawler support path.
The official policy links /crawler/ips.json but says that the placeholder does not currently publish crawler IP ranges. Therefore, do not create an allowlist or classify the registry token as genuine from the header or email address alone. Use Lyonl’s feedback path for confirmation of an observed unlisted identity.
Compare observed behavior with the canonical Lyonl search role without turning it into attribution. Public HTML, links, sitemaps, redirects, and conditional refreshes may be consistent with indexing. Lyonl states that HTML documents may be rendered with native Chromium, including SPA and hydration-heavy documents, within bounded byte, time, cache, and temporary-directory limits. Private endpoint access, form submission, account registration, high concurrency, repeated retries, or traffic outside verified evidence requires investigation. These observations establish impact, not operator identity or downstream use.
Evaluate /robots.txt independently. Confirm that it is served by the correct host, returns a successful status and text content type, and contains the exact group for the identity you intend to control. Lyonl uses separate tokens for web, image, and news crawling; do not apply lyonlbot rules to lyonl-crawler without confirmation. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate the unlisted token or secure private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If logs show the registry token, begin with a narrow report-only rule and preserve complete source evidence. Adapt the expression to your WAF provider:
{
"description": "Review unlisted Lyonl crawler candidates",
"expression": "lower(http.user_agent) contains \"lyonl-crawler\"",
"action": "log"
}
After Lyonl confirms the identity and you decide to restrict private routes, scope enforcement narrowly:
map $http_user_agent $block_lyonl_crawler_private {
default 0;
~*lyonl-crawler/0\.1 1;
}
server {
location ~ ^/(admin|private|account|internal|api)/ {
if ($block_lyonl_crawler_private) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof and must not be the sole authorization control. Do not invent IP ranges, reverse-DNS suffixes, or a Cloudflare verified-bot exception; the current policy says public verification is still in development. Test public HTML, SPA routes, redirects, sitemaps, conditional requests, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for the complete lyonl-crawler/0.1 header and preserve source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. If the token appears to be Lyonl traffic, report the IP, UTC timestamp, URL, full header, and sample log lines through https://lyonl.com/feedback as the policy requests for unlisted identities.
Check that your exact robots group expresses the intended public-search boundary. Test Crawl-delay only after the token is confirmed, and distinguish the unlisted token from lyonlbot, lyonlbot-image, and lyonlbot-news. Protect private routes with application authorization, not robots.txt.
Re-check the official Lyonl policy, feedback path, and /crawler/ips.json when public verification or IP ranges are published. Keep this profile at partially-documented until Lyonl adds lyonl-crawler to its official identity table or confirms it directly. Do not claim successful blocking or attribution from configuration alone; verify subsequent logs and operator response.
The current Lyonl policy documents a search-discovery service and explicitly says it does not intentionally bypass access controls, but those product statements must not be extended to a request that fails identity verification. Keep product-specific evidence separate from the registry token.
References
- LyonlBot crawler policy — official page listing the canonical web, image, and news identities, robots examples, search-discovery behavior, Chromium rendering, politeness, User-Agent spoofing warning, support path, and verification limitation; reviewed with the headless browser on 2026-08-25.
- Lyonl crawler policy URL — registry-linked URL; it redirected to the current bot policy during review.
- Lyonl crawler IP placeholder — policy-linked endpoint; the official page states it does not currently publish crawler IP ranges.
- Lyonl feedback — official path for reporting unlisted User-Agents and supplying IP/timestamp/path/log evidence.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.