← Bot Directory/FemtosearchBot
Bot directory / search-engine

FemtosearchBot: Robots.txt & Crawl Policy Reference

Technical reference for FemtosearchBot, the web crawler for the privacy-focused Femtosearch engine. Learn how to manage its access.

AI Summary: FemtosearchBot is the web crawler for Femtosearch, a privacy-focused search engine project developed by Grier Forensics. It crawls the web to build a search index and explicitly respects standard robots.txt directives. Blocking this bot will prevent your site from appearing in Femtosearch results.

Role and policy boundary

According to its official documentation, FemtosearchBot is a search engine crawler that discovers new and updated pages to add to the Femtosearch index. The crawler identifies itself using the User-Agent string Mozilla/5.0 (compatible; FemtosearchBot/1.0; http://femtosearch.com).

As a search engine crawler, its primary purpose is public indexing. The documentation explicitly states that it implements a policy of politeness by rate-limiting requests and respecting the robots exclusion protocol. Do not infer that its primary purpose is AI training. Note that the project's copyright dates to 2016, so crawl volume may be low compared to major engines.

If your logs confirm an exact FemtosearchBot token and you want to prevent your site from appearing in Femtosearch results, publish:

configuration / code
User-agent: FemtosearchBot
Disallow: /

For selective access (e.g., hiding internal sections from search):

configuration / code
User-agent: FemtosearchBot
Allow: /
Disallow: /private/
Disallow: /internal/

Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Because Femtosearch does not publish an official IP list or reverse DNS verification method, you must rely on behavioral analysis to verify the traffic.

Analyze behavior without assigning purpose prematurely. Requests for public pages, feeds, sitemaps, and metadata resemble authorized search indexing; deep traversal of private areas, high concurrency ignoring crawl delays, repeated retries, or private API access may indicate spoofing by malicious scrapers. These patterns demonstrate operational impact but cannot prove the operator or downstream use.

Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.

WAF and Nginx remediation examples

Once logs confirm an exact unwanted token (e.g., a spoofed bot ignoring robots.txt, or if you intentionally want to block Femtosearch globally), a narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:

configuration / code
{
  "description": "Block observed FemtosearchBot token",
  "expression": "lower(http.user_agent) contains \"femtosearchbot\"",
  "action": "block"
}

For Nginx, scope enforcement to private and high-cost routes while investigating public access:

configuration / code
map $http_user_agent $block_femtosearchbot {
    default 0;
    ~*FemtosearchBot 1;
}

server {
    location ~ ^/(private|internal|account|uploads|paywall|api)/ {
        if ($block_femtosearchbot) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and may block a legitimate client using the substring. Do not create an IP allowlist or broad network block without current operator evidence. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.

Review checklist

Search logs for every exact header that may be associated with FemtosearchBot and preserve representative requests. Record paths, response sizes, statuses, source networks, timing, and rate. Monitor for aggressive crawling behavior, as the lack of official IP verification makes this token susceptible to spoofing.

Decide whether your objective is to preserve search visibility, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group only for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately.

References

  1. Femtosearch Official Documentation — verified as active and documenting robots.txt compliance during the 2026-08-24 headless-browser review.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.