Bot directory / search-engine

MojeekBot: Robots.txt & Crawl Policy Reference

Technical reference for MojeekBot, including its published crawl frequency, robots matching behavior, meta-tag handling, forward-confirmed reverse DNS, and current IP range.

AI Summary: Mojeek documents MojeekBot as its web-search crawler. It should not request more than one page from a site within the same one-second period, does not support the nonstandard robots.txt Crawl-delay directive, follows the first matching MojeekBot record or the first * record, and obeys noindex, nocache, and nofollow meta tags. Mojeek publishes forward-confirmed reverse-DNS guidance and an IP range of 5.102.173.64/28; the User-Agent remains a signal that must be checked against source evidence.

Role and policy boundary

Mojeek’s official bot page identifies MojeekBot as the web crawler for the Mojeek search engine. It acknowledges that mistakes can happen and asks site owners to contact Mojeek if the bot crawls a page or directory it should not. The role described is search indexing; it does not grant access to private, authenticated, licensed, or origin-only material.

The registry records this versioned User-Agent:

configuration / code
MojeekBot/0.2 (archi; http://www.mojeek.com/bot.html)

The current Mojeek policy page documents MojeekBot behavior and provides a canonical source, while the registry string contains an archi label and an HTTP documentation link. Use the complete observed header in logs and do not authenticate a request from the name alone. A current page or future release may use a different version string, so keep the registry record and live header evidence distinct.

Mojeek states that it should not request more than one page from a site within the same one-second period, whether the response succeeds or fails. It explicitly says that MojeekBot does not support the nonstandard robots.txt Crawl-delay directive. To control crawling, Mojeek says the bot obeys the first robots record with a User-Agent containing MojeekBot; when no such record exists, it obeys the first record with User-Agent *.

A selective site-owner rule could be:

configuration / code
User-agent: MojeekBot
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /private/
Disallow: /account/
Disallow: /api/

For a complete exclusion:

configuration / code
User-agent: MojeekBot
Disallow: /

Do not add Crawl-delay expecting MojeekBot to interpret it. If a slower crawl is needed, use server-side rate limiting or contact Mojeek. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Mojeek’s official policy documents a forward-confirmed reverse-DNS procedure. For an observed IP, reverse-resolve it and confirm the hostname is within mojeek.com; then forward-resolve that hostname and confirm it returns the original IP. A mismatch means the request should not be treated as genuine MojeekBot traffic.

Mojeek’s worked example uses 5.102.173.71, reverse-resolving to crawl-5-102-173-71.mojeek.com, and then forward-resolving back to 5.102.173.71. The official IP JSON reviewed for this profile reports creation time 2025-05-14T15:44:06.000000 and prefix 5.102.173.64/28. Mojeek says the list may be updated at any time, so refresh it often rather than copying it as a permanent allowlist or blocklist.

Compare observed behavior with the documented search-indexing role. Public HTML, canonical metadata, feeds, and sitemaps may be consistent with indexing. Private endpoint access, high concurrency, repeated retries, unexpected downloads, or source IPs that fail DNS confirmation may indicate spoofing, abuse, a changed deployment, or another client. These observations establish impact, not operator identity or downstream use.

Evaluate /robots.txt independently. Confirm that the response is served by the correct host, returns a successful status and text content type, and places the intended MojeekBot record before any fallback * record where first-match semantics matter. Test disallow strings literally: Mojeek says a Disallow: /private string applies to URLs containing /private, including /private/ and /private.html.

Mojeek also documents page directives. A page containing:

configuration / code
<meta name="robots" content="noindex">

may still be retrieved, but Mojeek says it will not index the document or enter it into the search database. It also obeys nocache and nofollow. These directives do not authenticate the crawler or secure private routes; enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

If you want to recognize MojeekBot before enforcement, start with report-only logging and pair the header with DNS and IP checks. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe MojeekBot after DNS and range verification",
  "expression": "lower(http.user_agent) contains \"mojeekbot\"",
  "action": "log"
}

After verifying source evidence and deciding to restrict sensitive routes, use a narrow rule:

configuration / code
map $http_user_agent $block_mojeek_private {
    default 0;
    ~*MojeekBot/0\.2 1;
}

server {
    location ~ ^/(admin|private|account|internal|api)/ {
        if ($block_mojeek_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

The header match is not proof of origin and should not be used as the only authorization control. Refresh Mojeek’s official IP JSON and perform forward-confirmed reverse DNS before adding an allowlist. Do not use Crawl-delay as if Mojeek supports it; use an edge rate limiter, application controls, or the published contact path for load concerns.

Test public pages, URL paths containing disallow strings, feeds, sitemaps, redirects, noindex, nocache, nofollow, accounts, APIs, and approved integrations separately. Pair WAF controls with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search logs for the complete registered header and current MojeekBot versions. Record source IPs, ASNs, PTR results, forward lookups, paths, methods, response sizes, statuses, timing, and rate. Confirm the address falls within the current published range and that DNS is forward-confirmed before classifying the request as genuine.

Check the first applicable robots record and do not assume later groups override it. Confirm that path disallow strings match the URLs you intend to protect. Do not add Crawl-delay expecting MojeekBot to honor it; use server controls or contact Mojeek for a slower crawl.

Test whether noindex, nocache, and nofollow produce the intended indexing, cache, and link behavior, remembering that they do not stop retrieval or secure private content. Protect private routes with application authorization.

Refresh the official IP JSON frequently and re-check the bot page for User-Agent, DNS, frequency, and robots changes. Do not claim successful blocking from configuration alone; verify subsequent access logs and response behavior.

Keep the registry’s versioned header distinct from the current documentation. The official policy supports the MojeekBot role and behavior, but a spoofed or obsolete version string still requires source verification.

References

  1. MojeekBot policy — official documentation of purpose, frequency, robots first-record matching, disallow-string behavior, meta tags, DNS verification, and contact; reviewed with the headless browser on 2026-08-25.
  2. MojeekBot IP ranges — official JSON reporting creation time 2025-05-14T15:44:06.000000 and IPv4 prefix 5.102.173.64/28; reviewed with the headless browser on 2026-08-25.
  3. Mojeek contact — official contact path for crawler issues and enquiries.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.