← Bot Directory/search.marginalia.nu
Bot directory / search-engine

search.marginalia.nu: Robots.txt & Crawl Policy Reference

Technical reference for the Marginalia Search crawler, including its exact User-Agent, published IP range, crawl schedule, and independent discovery purpose.

AI Summary: Marginalia Search is an independent, open-source search engine prioritizing non-commercial content. Its official crawler documentation publishes the exact User-Agent and robots group search.marginalia.nu, a crawl interval of 8–10 weeks with daily RSS polling, and a verified IP range (193.183.0.162-193.183.0.174). The operator explicitly states the system uses "no AI." Verify the source IP against the published list before trusting the header.

Role and policy boundary

Marginalia Search describes itself as an independent, open-source internet search engine operating out of Sweden. Its published philosophy focuses on discovery and prioritizing non-commercial content, explicitly stating that it uses custom index/crawler software and "no AI."

The official crawler page documents this exact User-Agent:

configuration / code
search.marginalia.nu

The operator states that the crawler’s main refresh interval is 8–10 weeks, while the system also polls RSS feeds daily and performs live updates based on those feeds. The stated goal is to run an above-board, well-behaving search engine crawler.

The documentation specifies search.marginalia.nu as its robots.txt identifier. For selective access to public pages and RSS feeds:

configuration / code
User-agent: search.marginalia.nu
Allow: /public/
Allow: /blog/
Allow: /feed.xml
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

For a complete exclusion:

configuration / code
User-agent: search.marginalia.nu
Disallow: /

These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. The operator’s statement that it uses no AI means this crawler is not collecting data for model training or generative answers.

Layered verification

Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs. The official crawler page publishes an explicit IP range:

configuration / code
193.183.0.162-193.183.0.174

The operator also links to a full text list of IPs at https://search.marginalia.nu/crawler-ips.txt. To verify a request claiming to be Marginalia Search, confirm that the source IP falls within this documented range or the linked text file. A User-Agent match from an unlisted IP is a spoof and should not be treated as authentic Marginalia traffic.

Compare observed traffic with the documented search-indexing and RSS-polling roles. Public HTML, metadata, RSS feeds, sitemaps, and ordinary assets are consistent with discovery. Private endpoints, account routes, APIs, high concurrency, or unexpected bulk downloads establish impact and load risk, not permission. The stated 8–10 week interval and daily RSS polling provide a baseline for expected traffic patterns.

Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact search.marginalia.nu group you intend to publish. Check precedence, path matching, and actual request behavior. Meta directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

Begin in report-only mode and correlate the exact User-Agent with the published IP range. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe search.marginalia.nu candidates before enforcement",
  "expression": "http.user_agent eq \"search.marginalia.nu\"",
  "action": "log"
}

For a deliberately restricted private route, use the exact token only after verifying the source IP against the published list:

configuration / code
map $http_user_agent $block_marginalia_private {
    default 0;
    "search.marginalia.nu" 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_marginalia_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

This rule is route-scoped and does not prove identity. Do not hard-code the reviewed IP range indefinitely; fetch the operator’s text list if you implement a network-level allowlist. Do not invent additional ASN exceptions, rates, or training permissions.

A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action. The operator provides contact@marginalia-search.com for complaints if the crawler misbehaves.

Test public pages, RSS feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact search.marginalia.nu string, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Verify the source IP against the operator’s published range (193.183.0.162-193.183.0.174) or the linked IP text file. Treat requests from outside this range as spoofed.

Check your robots file for an exact search.marginalia.nu group. Test precedence and path behavior, and verify actual logs after publishing. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.

Review search indexing, AI input, reference use, and model-training decisions separately. The operator explicitly states the engine uses no AI and focuses on non-commercial discovery.

Keep the profile at documented: the operator publishes the exact User-Agent, robots identifier, IP range, crawl schedule, and contact email. Re-check the crawler documentation and IP list when the source network or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.

References

  1. Marginalia Search Crawler Information — official page documenting the User-Agent, robots identifier, IP range, crawl interval, and contact email; reviewed with the headless browser on 2026-08-25.
  2. Marginalia Search About — official context describing the independent, non-commercial, "no AI" search philosophy.
  3. Marginalia Crawler IPs — linked text file for automated IP verification.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.