Bot directory / search-engine

FleebsBot: Robots.txt & Crawl Policy Reference

Technical reference for FleebsBot, the web crawler for the Fleebs real-time search engine. Learn how to manage its access to your site.

AI Summary: FleebsBot is the official web crawler for Fleebs, a real-time search engine. It crawls the web to discover and analyze new information for its search index. According to its official documentation, it respects standard robots.txt directives. Blocking this bot will prevent your site from appearing in Fleebs search results.

Role and policy boundary

FleebsBot operates on behalf of the fleebs.com real-time search engine. Its official documentation identifies two main variants of the crawler: one for discovering new information (identifying with FleebsBot/1~w) and another for analyzing found web pages (identifying with FleebsBot/1~c).

As a search engine crawler, its primary purpose is public indexing to deliver news and real-time updates to its users. Blocking FleebsBot means your content will not be discoverable on their platform. Do not infer that its primary purpose is AI training, though indexed data may inform their search algorithms.

If your logs confirm an exact FleebsBot token and you want to prevent your site from appearing in Fleebs search results, publish:

configuration / code
User-agent: FleebsBot
Disallow: /

For selective access (e.g., hiding private sections from search):

configuration / code
User-agent: FleebsBot
Allow: /
Disallow: /private/
Disallow: /internal/
Disallow: /api/

Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Fleebs does not currently publish a canonical list of IP addresses or a reverse DNS verification method on its main documentation page.

Analyze behavior without assigning purpose prematurely. Requests for public pages, feeds, sitemaps, and metadata resemble authorized search indexing; deep traversal of private areas, high concurrency ignoring crawl delays, repeated retries, or private API access may indicate spoofing by malicious scrapers. These patterns demonstrate operational impact but cannot prove the operator or downstream use.

Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.

WAF and Nginx remediation examples

Once logs confirm an exact unwanted token (e.g., a spoofed bot ignoring robots.txt, or if you intentionally want to block Fleebs globally), a narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:

configuration / code
{
  "description": "Block observed FleebsBot token",
  "expression": "lower(http.user_agent) contains \"fleebsbot\"",
  "action": "block"
}

For Nginx, scope enforcement to private and high-cost routes while investigating public access:

configuration / code
map $http_user_agent $block_fleebsbot {
    default 0;
    ~*FleebsBot 1;
}

server {
    location ~ ^/(private|internal|account|uploads|paywall|api)/ {
        if ($block_fleebsbot) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and may block a legitimate client using the substring. Do not create an IP allowlist or broad network block without current operator evidence. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.

Review checklist

Search logs for every exact header that may be associated with FleebsBot and preserve representative requests. Record paths, response sizes, statuses, source networks, timing, and rate. Monitor for aggressive crawling behavior, as the lack of official IP verification makes this token susceptible to spoofing.

Decide whether your objective is to preserve search visibility, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group only for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately.

References

  1. FleebsBot Official Documentation — verified as active and documenting robots.txt compliance during the 2026-08-24 headless-browser review.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.