Bot directory / scraper

FediIndex: Robots.txt & Crawl Policy Reference

Technical reference for FediIndex. Learn how to investigate Fediverse server indexing traffic when the documentation is partially inaccessible.

AI Summary: FediIndex is a community-maintained web crawler associated with indexing the Fediverse (such as Mastodon and Lemmy instances). The registry-linked policy URL (fedi.wrm.sr/about) blocked our automated review, preventing public verification of its exact IP ranges and robots behavior. It historically identifies as FediIndex/1.0 (+https://fedi.wrm.sr/about). Treat it as community discovery traffic, but use layered controls to verify its identity.

Role and policy boundary

The registry describes FediIndex as a Fediverse server indexing and discovery service web crawler. While the domain fedi.wrm.sr is active, it actively blocked our headless-browser review, meaning the full policy cannot be publicly verified programmatically.

Based on its name and community context, this crawler likely discovers and indexes decentralized social media instances (the Fediverse) to map the network or provide directory services. Do not infer that it collects general AI-training data or builds a commercial public search index outside of the Fediverse ecosystem.

If your logs confirm an exact FediIndex token and you want to communicate a restriction (e.g., if you run a private Fediverse instance), publish:

configuration / code
User-agent: FediIndex
Disallow: /

For selective access:

configuration / code
User-agent: FediIndex
Allow: /about/
Allow: /api/v1/instance
Disallow: /private/
Disallow: /internal/

Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Because the primary documentation was inaccessible, do not classify traffic from a partial substring alone.

Analyze behavior without assigning purpose prematurely. Requests for public instance metadata (like /api/v1/instance on Mastodon) or public timelines may resemble authorized Fediverse discovery; deep traversal of private user profiles, high concurrency, repeated retries, or aggressive API polling may indicate misconfiguration or abuse. These patterns demonstrate operational impact but cannot prove the operator or downstream use.

Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.

WAF and Nginx remediation examples

Once logs confirm an exact unwanted token, a narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:

configuration / code
{
  "description": "Block observed FediIndex token",
  "expression": "lower(http.user_agent) contains \"fediindex\"",
  "action": "block"
}

For Nginx, scope enforcement to private and high-cost routes while investigating public access:

configuration / code
map $http_user_agent $block_fediindex {
    default 0;
    ~*FediIndex 1;
}

server {
    location ~ ^/(private|internal|account|uploads|paywall)/ {
        if ($block_fediindex) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and may block a legitimate client using the substring. Do not create an IP allowlist or broad network block without current operator evidence. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.

Review checklist

Search logs for every exact header that may be associated with FediIndex and preserve representative requests. Record paths, response sizes, statuses, source networks, timing, and rate. Re-check the registry-linked policy URL before upgrading this profile; during this review the documentation page blocked automated access.

Decide whether your objective is to preserve Fediverse discoverability, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group only for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately. Revisit the profile if an accessible live policy becomes available.

References

  1. Registry-linked FediIndex policy URL — blocked automated access during the 2026-08-24 headless-browser review.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.