← Bot Directory/FyndSearchEngine-ReCrawler
Bot directory / scraper

FyndSearchEngine-ReCrawler: Robots.txt & Crawl Policy Reference

Technical reference for FyndSearchEngine-ReCrawler, a partially documented re-indexing label associated with the active fynd.bot search service.

AI Summary: FyndSearchEngine-ReCrawler is a partially documented crawler label associated with the active fynd.bot web search service. The public homepage confirms a search-engine role, but its /bot path returns 404 and does not publish re-indexing policy, IP ranges, rate limits, or verification instructions. Treat the registry User-Agent as an unverified signal until it appears in your logs, then use layered controls.

Role and policy boundary

The registry describes FyndSearchEngine-ReCrawler as a Fynd search-engine re-indexing crawler and records the identity as FyndSearchEngine-ReCrawler (by fynd.bot; https://fynd.bot). Headless-browser review confirms that fynd.bot currently presents a web search service titled “fynd — Search the World Wide Web.” The linked /bot path returns 404, so the service’s exact re-crawl rules and policy boundary remain undocumented.

The defensible purpose classification is search indexing and re-indexing: the public site is a search engine, and the registry description states re-indexing. That does not prove that every request using this label is genuine, nor does it prove that the crawler collects AI-training data, stores full page content, or follows a particular revisit schedule. A re-crawl label describes a registry category, not verified operational behavior.

If your logs confirm the exact token and you want to exclude it from search discovery, publish:

configuration / code
User-agent: FyndSearchEngine-ReCrawler
Disallow: /

For selective access to public reference material while protecting sensitive areas:

configuration / code
User-agent: FyndSearchEngine-ReCrawler
Allow: /public-reference/
Allow: /docs/
Disallow: /private/
Disallow: /internal/
Disallow: /api/

Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Begin with raw access logs. Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. The accessible Fynd pages do not publish a canonical IP list or a reverse-DNS procedure, so do not treat the User-Agent alone as authentication.

Compare observed behavior with a re-indexing hypothesis without turning that hypothesis into a fact. Repeated requests for public HTML, canonical metadata, feeds, sitemaps, or pages previously observed may be consistent with index refreshes. High-concurrency traversal, repeated retries, large asset downloads, private endpoint access, or disregard for an explicit site policy may indicate spoofing, abuse, or a separate client. These observations establish impact, not operator identity or downstream use.

Evaluate /robots.txt independently. Confirm that the response is served from the correct host, has a successful status and text content type, contains an exact FyndSearchEngine-ReCrawler group, and matches the paths you intend to restrict. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not secure private routes and do not prove that an undocumented crawler has accepted an opt-out. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

After logs confirm the exact token and your policy decision is to restrict it, a narrow WAF rule can match the declared identity. Replace the expression with the syntax and field names of your provider:

configuration / code
{
  "description": "Block observed FyndSearchEngine-ReCrawler token",
  "expression": "lower(http.user_agent) contains \"fyndsearchengine-recrawler\"",
  "action": "block"
}

For Nginx, scope the first control to private and expensive routes while keeping public search access under observation:

configuration / code
map $http_user_agent $block_fynd_recrawler {
    default 0;
    ~*FyndSearchEngine-ReCrawler 1;
}

server {
    location ~ ^/(private|internal|account|uploads|paywall|api)/ {
        if ($block_fynd_recrawler) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof or evade and can also catch a legitimate integration if the expression is too broad. Do not invent an IP allowlist because none is published in the reviewed source. Start with a report-only rule, test representative pages and integrations, then enforce narrowly. Pair edge rules with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search the logs for the complete registered header, including the by fynd.bot and documentation URL text where present. Record source IPs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Re-check https://fynd.bot/ and the missing /bot path when new traffic appears; upgrade this profile only if the operator publishes a current re-crawling policy, canonical UA, source verification method, or robots guidance.

Decide whether your objective is to preserve Fynd search visibility, prevent content extraction, protect private material, or reduce crawl load. Publish an exact robots group only if that is the intended search-visibility trade-off, enforce private routes with WAF and application controls, and test HTML, feeds, sitemaps, media, uploads, and APIs separately. Do not claim successful blocking from a configuration change alone; verify the result in subsequent access logs.

References

  1. Fynd search homepage — active public search service observed during the 2026-08-24 headless-browser review.
  2. Fynd crawler path — returned 404 Not Found during the same review; no re-crawling policy was inferred from the missing page.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.