Greppr Web Crawler: Robots.txt & Crawl Policy Reference
Technical reference for the Greppr search crawler, including its public search-index role, observed robots.txt rule, and limits around User-Agent and source verification.
AI Summary: Greppr is an active search service presenting a human-scale web index, and its public robots.txt disallows
/searchfor all User-Agents. The file does not publish a Greppr-specific group, IP range, reverse-DNS verification method, or Crawl-delay. Treat the registry User-Agent as a useful log-search signal, verify traffic independently, and keep public search visibility decisions separate from private access control.
Role and policy boundary
Greppr’s official homepage describes a live web index focused on clarity, free from AI noise, ads, and tracking. During review it displayed version 0.21.1 BETA and stated that its collection contained 47,931,092 essential pages. These observations support a search-indexing role for the service, but the homepage does not by itself authenticate every request carrying the registry header.
The inventory records this User-Agent:
Mozilla/5.0 (compatible; Greppr Web Crawler/0.12.0 (https://greppr.org/))
Greppr’s current robots file contains:
User-agent: *
Disallow: /search
There is no Greppr-specific group in the reviewed file. Therefore, do not claim that Greppr has published a dedicated opt-out, Crawl-delay, IP range, or reverse-DNS procedure. The global /search rule is an observation about the current file, not proof of how a particular crawler handles every other path.
If you want to exclude the observed Greppr token from public discovery, a site-owner rule could be:
User-agent: Greppr Web Crawler
Disallow: /
For selective access to approved public pages while preserving the global restriction:
User-agent: Greppr Web Crawler
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
These are defensive examples, not a recovered Greppr policy. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Begin with access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Greppr’s reviewed robots file publishes no source network or DNS verification method. Do not treat the header alone as authentication.
Compare observed behavior with a search-indexing hypothesis without turning it into attribution. Requests for public HTML, canonical metadata, feeds, sitemaps, and ordinary assets may be consistent with an indexer. High-concurrency traversal, private endpoint access, repeated retries, unusual downloads, or traffic that ignores your policy may indicate spoofing, abuse, a fork, or a different client. These observations establish impact, not operator identity or downstream use.
Evaluate /robots.txt independently. Confirm the response is served by the correct host, has a successful status and text content type, and contains the exact global and bot-specific groups you intend to publish. Test /search, public pages, private routes, and APIs separately; the observed global rule must not be generalized to paths it does not mention. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate Greppr or secure private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If logs show a repeatable unwanted Greppr token, begin with report-only mode and correlate the header with source evidence. Adapt the expression to your WAF provider:
{
"description": "Observe Greppr Web Crawler requests",
"expression": "lower(http.user_agent) contains \"greppr web crawler\"",
"action": "log"
}
After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:
map $http_user_agent $block_greppr_crawler {
default 0;
~*Greppr[ ]Web[ ]Crawler 1;
}
server {
location ~ ^/(admin|account|private|internal|api)/ {
if ($block_greppr_crawler) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof and can catch an authorized search integration if it is too broad. Do not invent an IP allowlist or reverse-DNS suffix from the homepage or robots file. Test public pages, /search, feeds, sitemaps, media, account flows, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for the complete registered header and preserve representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Confirm whether requests match public indexing behavior and whether the source is independently attributable; do not classify them as authentic from the token alone.
Re-check Greppr’s homepage and robots.txt whenever its beta version or index behavior changes. Keep the profile at partially-documented until the operator publishes a canonical crawler page, source-verification method, IP ranges, rate guidance, or a Greppr-specific robots group. Treat the current global /search rule as limited evidence about the present file, not a guarantee about future behavior.
Decide whether the objective is to preserve Greppr visibility, restrict search indexing, protect private material, or reduce crawl load. Publish exact robots rules, enforce private routes with application and WAF controls, and do not claim successful blocking from configuration alone; verify subsequent logs and response behavior.
References
- Greppr homepage — active search service page observed with version
0.21.1 BETAand the stated web-index description, reviewed with the headless browser on 2026-08-25. - Greppr robots.txt — current file observed with
User-agent: *andDisallow: /search; no Greppr-specific group or source-verification data was present during review. - Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.