← Bot Directory/WepchSearchEngine
Bot directory / search-engine

WepchSearchEngine: Robots.txt & Crawl Policy Reference

Technical reference for WepchSearchEngine, a crawler for an independent search engine project, including its robots.txt compliance and Crawl-delay support.

AI Summary: WepchSearchEngine is a web crawler for an independent, privacy-focused search engine project by Wepch. The official documentation confirms it respects the WepchSearchEngine robots.txt token and supports the Crawl-delay directive. The operator notes the project is currently delayed due to financial difficulties. Because the operator does not publish a static IP list or a specific reverse-DNS verification method, use standard network validation and behavioral analysis before trusting the header.

Role and policy boundary

Wepch operates WepchSearchEngine to crawl the web and build an independent search index. The registry records this User-Agent string:

configuration / code
Mozilla/5.0 (compatible; WepchSearchEngine; +https://www.wepch.com/)

The operator explicitly states that WepchSearchEngine respects the standard robots.txt protocol and identifies itself using the WepchSearchEngine token. It also explicitly supports the Crawl-delay directive to manage server load. For selective access to public pages with a rate limit:

configuration / code
User-agent: WepchSearchEngine
Crawl-delay: 30
Allow: /public/
Allow: /articles/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

For a complete exclusion from the Wepch index:

configuration / code
User-agent: WepchSearchEngine
Disallow: /

These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. The search-indexing purpose does not establish permission for generative AI answers or third-party model training.

Layered verification

Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs. The reviewed Wepch documentation does not publish a static IP list or a specific reverse-DNS verification procedure.

Because a User-Agent is easily spoofed, do not treat the header alone as authentication. Check source IP and DNS evidence independently. A WepchSearchEngine header from a residential network, a known proxy, or a cloud provider without an established Wepch affiliation should not be treated as authentic traffic.

Compare observed traffic with the documented search-indexing role. Public HTML, metadata, feeds, sitemaps, and ordinary assets are consistent with discovery. Private endpoints, account routes, APIs, licensed content, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not permission.

Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact WepchSearchEngine group you intend to publish. Check precedence, path matching, and actual request behavior. Meta directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

Begin in report-only mode and correlate the WepchSearchEngine User-Agent with independent network verification. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe WepchSearchEngine candidates before enforcement",
  "expression": "lower(http.user_agent) contains \"wepchsearchengine\"",
  "action": "log"
}

For a deliberately restricted private route, use the exact token only after verifying the source network:

configuration / code
map $http_user_agent $block_wepch_private {
    default 0;
    ~*WepchSearchEngine 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_wepch_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

This rule is route-scoped and does not prove identity. Do not invent an IP allowlist, Wepch ASN exception, rate, or training permission without operator documentation.

A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action. The operator provides nokrogi@gmail.com for questions or bug reports.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact WepchSearchEngine string, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Verify the source network independently before trusting the client. Treat requests from unverified networks as spoofed.

Check your robots file for an exact WepchSearchEngine group. Test precedence and path behavior, and verify actual logs after publishing. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.

Review search indexing, AI input, reference use, and model-training decisions separately. The operator explicitly states the crawler is used for its independent search engine.

Keep the profile at documented-limit: the operator publishes the crawler identity, robots compliance, and purpose, but lacks a static IP list or reverse-DNS verification method, and the project is currently delayed. Re-check the crawler documentation when the source network or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.

References

  1. Wepch Crawler Information — official page documenting the WepchSearchEngine robots.txt token, Crawl-delay support, and project status; reviewed with the headless browser on 2026-08-25.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.