SeznamBot: Robots.txt & Crawl Policy Reference
Technical reference for SeznamBot, the Seznam.cz search engine crawler, including its robots.txt compliance, Request-rate extension, and indexing controls.
AI Summary: SeznamBot is the official crawler for the Czech search engine Seznam.cz. Its documentation confirms full compliance with robots.txt and supports the nonstandard
Request-ratedirective to control crawl frequency (down to 1 request per 10 seconds) and time-of-day limits. It also obeysnoindexandnofollowmeta tags. The registry provides a historical3.2-test1-1version; preserve the complete observed User-Agent and verify the source IP against Seznam’s published JSON list before trusting the header.
Role and policy boundary
SeznamBot is the web crawler for Seznam.cz’s full-text search engine. The operator’s official Help Search pages document its behavior, explaining that it discovers and downloads content to build the search index.
The registry records this historical User-Agent pattern:
Mozilla/5.0 (compatible; SeznamBot/3.2-test1-1; +http://napoveda.seznam.cz/en/seznambot-intro/)
The exact version and documentation URL in the header may change over time (e.g., pointing to o-seznam.cz), but the SeznamBot identifier remains stable.
The operator explicitly states that SeznamBot fully complies with the robots exclusion standard. For selective access to public pages:
User-agent: SeznamBot
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
For a complete exclusion, the operator provides this example:
User-agent: SeznamBot
Disallow: /
SeznamBot also recognizes the nonstandard Request-rate directive to control crawl frequency. The minimum rate is 1 document every 10 seconds. You can also specify time-of-day windows (in UTC):
User-agent: SeznamBot
Request-rate: 1/10s
Request-rate: 400/1h 1800-1900
These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. The search-indexing purpose does not establish permission for generative AI answers or third-party model training.
Layered verification
Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs. Seznam publishes a JSON list of its crawler IPs to help webmasters verify traffic.
To verify a request claiming to be SeznamBot, confirm that the source IP is present in the official Seznam JSON list or resolves via reverse DNS to a Seznam.cz hostname (with a matching forward lookup). A User-Agent match from an unlisted or unverified IP is a spoof and should not be treated as authentic Seznam traffic.
Compare observed traffic with the documented search-indexing role. Public HTML, metadata, links, feeds, sitemaps, and ordinary assets are consistent with discovery. Private endpoints, account routes, APIs, licensed content, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not permission.
Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact SeznamBot group you intend to publish. Seznam explicitly supports page-level indexing controls via HTML meta tags:
<meta name="robots" content="noindex, nofollow">
It also supports the rel="nofollow" attribute on individual links to prevent the crawler from following them. These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
Begin in report-only mode and correlate the SeznamBot User-Agent with Seznam’s IP list or DNS verification. Adapt the expression to your WAF provider:
{
"description": "Observe SeznamBot candidates before enforcement",
"expression": "lower(http.user_agent) contains \"seznambot\"",
"action": "log"
}
For a deliberately restricted private route, use the exact token only after verifying the source IP against the published list:
map $http_user_agent $block_seznam_private {
default 0;
~*SeznamBot 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_seznam_private) { return 403; }
try_files $uri $uri/ =404;
}
}
This rule is route-scoped and does not prove identity. Fetch the operator’s JSON list dynamically if you implement a network-level allowlist. Do not invent additional ASN exceptions, rates, or training permissions.
A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.
Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.
Review checklist
Search logs for the exact SeznamBot string, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.
Verify the source IP against Seznam’s published JSON list or perform reverse/forward DNS verification confirming a Seznam.cz hostname. Treat requests from outside this list as spoofed.
Check your robots file for an exact SeznamBot group. Test Request-rate precedence, path behavior, and verify actual logs after publishing. Use META tags or X-Robots-Tag for indexing preferences and authentication for private resources.
Review search indexing, AI input, reference use, and model-training decisions separately. The operator explicitly states the engine is a full-text search index.
Keep the profile at documented: the operator publishes the exact User-Agent identifier, robots compliance, Request-rate extension, meta tag support, and IP verification. Re-check the crawler documentation and IP list when the source network or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.
References
- Seznam.cz Crawling Control — official page documenting SeznamBot's robots.txt compliance and Request-rate directive; reviewed with the headless browser on 2026-08-25.
- Seznam.cz Indexing Control — official page documenting SeznamBot's support for noindex, nofollow, and rel="nofollow".
- Seznam.cz SeznamBot Crawler — official page providing crawler identity and IP verification JSON link.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.