lyonl-asset-proxy: Robots.txt & Crawl Policy Reference
Technical reference for the lyonl-asset-proxy registry label, with current Lyonl crawler documentation and explicit limits around its unlisted identity.
AI Summary: The registry records
lyonl-asset-proxy/1.0 (+https://lyonl.com/crawler), but Lyonl’s current first-party crawler policy lists onlyLyonlBot,LyonlBot Image, andLyonlBot News; the asset-proxy token is not listed. Lyonl documents separate robots tokens, User-Agent spoofing risk, and a verification method still in development. Treat this token as product-specific and unverified until Lyonl confirms it.
Role and policy boundary
The inventory describes lyonl-asset-proxy as a Lyonl search-engine asset proxy and records this User-Agent:
lyonl-asset-proxy/1.0 (+https://lyonl.com/crawler)
The registry documentation URL redirected to Lyonl’s current https://lyonl.com/bot.html policy. That first-party page describes Lyonl Search and three listed crawler identities: canonical LyonlBot, specialized LyonlBot Image, and specialized LyonlBot News. It does not list lyonl-asset-proxy in the identity table. The page says that an unlisted Lyonl User-Agent should be reported with its IP, timestamp, requested URL, and full header.
The closest documented asset-related identity is LyonlBot Image, whose stated purpose is image discovery, rendered-document image extraction, and asset-rendering support. Do not silently merge the registry token with that identity. The asset-proxy role remains a registry hypothesis until Lyonl confirms the exact token and behavior.
Lyonl’s policy says the canonical crawler visits public URLs, does not intentionally bypass authentication, paywalls, CAPTCHAs, or access controls, and does not submit forms, register accounts, or perform transactions. Those statements apply to the listed Lyonl crawler service and should not be extended to an unlisted token or a spoofed request.
Because the exact token is not listed, do not publish a rule that assumes lyonl-asset-proxy is an official Lyonl identity. If logs and Lyonl support confirm it, use an exact rule:
User-agent: lyonl-asset-proxy
Disallow: /
For selective access after confirmation:
User-agent: lyonl-asset-proxy
Allow: /public/
Allow: /images/public/
Disallow: /private/
Disallow: /original-assets/
Disallow: /admin/
Disallow: /api/
Crawl-delay: 10
These are site-owner examples, not a recovered asset-proxy policy. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Lyonl’s policy explicitly warns that User-Agent strings can be spoofed and says a public verification method is still being developed. Until IP ranges are published, it asks site owners to report the IP, UTC timestamp, URL, full header, and relevant log lines through its crawler support path.
The linked /crawler/ips.json placeholder does not currently publish crawler IP ranges. Therefore, do not create an IP allowlist or claim that a request is genuine from the registry token. Send representative evidence to Lyonl’s feedback path if you need identity confirmation or a policy decision.
Compare observed behavior with an asset or public-document discovery hypothesis without turning it into attribution. Public image assets, rendered documents, HTML, and search-index resources may be consistent with the documented Lyonl family, but private originals, account routes, APIs, and high-volume downloads require investigation. These observations establish impact and data risk, not operator identity or downstream use.
Evaluate /robots.txt independently. Confirm that the response is served by the correct host, returns a successful status and text content type, and contains an exact rule for the identity you intend to control. Lyonl says separate tokens allow site owners to control web, image, and news crawlers precisely; do not apply an image rule to the unlisted asset-proxy token without confirmation. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, noimageindex">
X-Robots-Tag: noindex, noimageindex
These signals do not authenticate the proxy or secure private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If logs show the registry token, start in report-only mode and include complete source evidence. Adapt the expression to your WAF provider; it identifies an observed header only:
{
"description": "Review unlisted Lyonl asset-proxy candidates",
"expression": "lower(http.user_agent) contains \"lyonl-asset-proxy\"",
"action": "log"
}
After Lyonl confirms the identity and you decide to restrict private assets, scope enforcement to sensitive routes:
map $http_user_agent $block_lyonl_asset_private {
default 0;
~*lyonl-asset-proxy/1\.0 1;
}
server {
location ~ ^/(private|original-assets|account|internal|api)/ {
if ($block_lyonl_asset_private) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof and must not be the sole authorization control. Do not invent IP ranges, reverse-DNS suffixes, or a Cloudflare verified-bot exception; the policy says public verification is still in development. Test public images, rendered documents, original assets, HTML, sitemaps, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, hotlink controls, and anomaly detection.
Review checklist
Search logs for the complete lyonl-asset-proxy header and preserve source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Report an unlisted Lyonl token through https://lyonl.com/feedback with the IP, UTC timestamp, URL, full header, and relevant log lines when identity confirmation is needed.
Check that your exact robots group expresses the intended public asset and private-origin boundary. Test Crawl-delay only after the token is confirmed and verify whether separate lyonlbot-image or lyonlbot-news rules are the actual identities seen in logs. Protect original assets, accounts, and APIs with application authorization, not robots.txt.
Re-check the official Lyonl policy, support path, and /crawler/ips.json placeholder when the service publishes a verification method or ranges. Keep this profile at partially-documented until Lyonl adds lyonl-asset-proxy to its official table or confirms it directly. Do not claim successful blocking or attribution from configuration alone; verify subsequent logs and operator response.
The current Lyonl policy describes search discovery and image/news specialization, but it does not establish training use for the unlisted asset-proxy token. Keep product-specific statements separate from unverified downstream behavior.
References
- LyonlBot crawler policy — official page listing the canonical, image, and news crawler identities, robots examples, public-document behavior, User-Agent spoofing warning, support path, and verification limitation; reviewed with the headless browser on 2026-08-25.
- Lyonl crawler IP placeholder — linked official endpoint; policy text states that it does not currently publish crawler IP ranges, and the browser capture exposed no range.
- Lyonl feedback — official support path for reporting unlisted User-Agents and supplying IP/timestamp/path/log evidence.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.