Bot directory / search-engine

SeekportBot: Robots.txt & Crawl Policy Reference

Technical reference for SeekportBot, the search engine crawler operated by SISTRIX, including its exact User-Agent, published IP list, and robots controls.

AI Summary: SeekportBot is a search engine crawler operated by the German platform intelligence provider SISTRIX. Its official documentation publishes the exact User-Agent Mozilla/5.0 (compatible; SeekportBot; +https://bot.seekport.com), states that it obeys robots rules and Crawl-delay, and provides a complete IP list at /seekportbot_ips.txt. The operator states the search engine does not profile users or serve advertising. Verify the source IP against the published list before trusting the header.

Role and policy boundary

Seekport’s official bot page describes the service as a public, free, and independent internet search engine originally founded in 2003 and operated by SISTRIX since December 2014. The page states that Seekport does not store user data, does not profile users, and operates without advertising. To update its search index, SeekportBot scours the public internet.

The documentation publishes this exact User-Agent:

configuration / code
Mozilla/5.0 (compatible; SeekportBot; +https://bot.seekport.com)

The operator explicitly states that SeekportBot follows Disallow instructions in robots.txt and obeys the Crawl-delay directive. The documentation provides this example for slowing down the crawler:

configuration / code
User-agent: SeekportBot
Crawl-delay: 2

For selective access to public pages:

configuration / code
User-agent: SeekportBot
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

For a complete exclusion, the operator provides this example:

configuration / code
User-agent: SeekportBot
Disallow: /

These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. The search-indexing purpose does not establish permission for generative AI answers or third-party model training.

Layered verification

Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs. The official crawler page provides a complete and up-to-date list of IP addresses in use by SeekportBot at:

configuration / code
https://bot.seekport.com/seekportbot_ips.txt

To verify a request claiming to be SeekportBot, confirm that the source IP is present in this published list. A User-Agent match from an unlisted IP is a spoof and should not be treated as authentic SISTRIX traffic.

Compare observed traffic with the documented search-indexing role. Public HTML, metadata, links, feeds, sitemaps, and ordinary assets are consistent with discovery. Private endpoints, account routes, APIs, licensed content, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not permission. The operator states that SeekportBot tries not to overload websites and makes as few calls per second as possible.

Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact SeekportBot group you intend to publish. Check precedence, path matching, and actual request behavior. Meta directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

Begin in report-only mode and correlate the exact User-Agent with the published IP list. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe SeekportBot candidates before enforcement",
  "expression": "http.user_agent eq \"Mozilla/5.0 (compatible; SeekportBot; +https://bot.seekport.com)\"",
  "action": "log"
}

For a deliberately restricted private route, use the exact token only after verifying the source IP against the published list:

configuration / code
map $http_user_agent $block_seekport_private {
    default 0;
    "Mozilla/5.0 (compatible; SeekportBot; +https://bot.seekport.com)" 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_seekport_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

This rule is route-scoped and does not prove identity. Fetch the operator’s text list dynamically if you implement a network-level allowlist. Do not invent additional ASN exceptions, rates, or training permissions.

A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action. The operator provides seekportbot@sistrix.com for questions or comments if the crawler misbehaves.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact SeekportBot string, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Verify the source IP against the operator’s published list at https://bot.seekport.com/seekportbot_ips.txt. Treat requests from outside this list as spoofed.

Check your robots file for an exact SeekportBot group. Test Crawl-delay precedence, path behavior, and verify actual logs after publishing. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.

Review search indexing, AI input, reference use, and model-training decisions separately. The operator explicitly states the engine is an independent search alternative and does not profile users.

Keep the profile at documented: the operator publishes the exact User-Agent, robots identifier, IP list, crawl behavior, and contact email. Re-check the crawler documentation and IP list when the source network or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.

References

  1. Seekport Bot — official page documenting the User-Agent, robots identifier, Crawl-delay support, operator (SISTRIX), and contact email; reviewed with the headless browser on 2026-08-25.
  2. SeekportBot IP list — official text file containing the complete list of IPs in use.
  3. Seekport — search engine homepage.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.