7Siters: Robots.txt & Crawl Policy Reference
Technical reference for the 7Siters website-statistics crawler. Learn its first-party opt-out file, Apache blocking example, and verification limits.
AI Summary: 7Siters is a website-statistics and indexing service whose first-party 7ooo.ru page describes collecting statistics about sites and internal pages. Its webmaster section gives a site-removal workflow using a root-level
7Siters.txtfile containingdelete, plus an Apache.htaccessUser-Agent block. The page does not publish source IP ranges or a formal robots.txt policy, so verify requests locally and treat the opt-out mechanism as the primary documented control.
Role and policy boundary
The 7Siters page describes research into the Internet and its segments, including collecting independent statistics about the number of sites and internal pages. It says collected data is analyzed and published through the service. This supports a website-statistics and discovery purpose, but it does not prove AI training, search indexing for a third party, or access to private content.
The page identifies the crawler in its blocking example as 7Siters; the registry provides 7Siters/1.2 ( https://7ooo.ru/siters/ ). The first-party page also provides support@7siters.com for questions and proposals. No source IP ranges, crawl frequency, or formal robots.txt behavior was found in the reviewed webmaster section. Do not treat the presence of a User-Agent as authentication.
The documented opt-out method is a root-level file named 7Siters.txt containing delete. The page also provides a site-removal form and says the domain should be entered without http:// or https://. For sites that prefer a server-side block, it provides this Apache pattern:
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} ^.*7Siters.* [NC,OR]
RewriteRule ^(.*)$ - [F,L]
A robots communication group can complement those controls:
User-agent: 7Siters
Disallow: /
The first-party opt-out instructions are more specific than the robots group. Robots.txt is advisory and cannot protect private content; use authentication and authorization for sensitive routes.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Compare the observed header with the registry value but remember that a client can copy it.
Check whether the request pattern targets public pages, internal links, feeds, sitemaps, media, or private paths. The site's description supports a statistics and collection context, but behavior still cannot prove the downstream use of a particular URL or asset. Requests for login, admin, account, or API endpoints should be treated as an access-control issue regardless of the bot's stated purpose.
Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact User-agent: 7Siters group, and path match. If you use the documented 7Siters.txt opt-out, verify its exact filename, root placement, content, response status, and whether the operator has removed the domain from its database. These are separate controls and should not be conflated.
Page-level metadata may express a discoverability preference:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives do not secure private routes and do not replace the first-party deletion mechanism or application authorization.
WAF and Nginx remediation examples
If the documented opt-out has not stopped observed traffic, or if you need an immediate local block, use a narrow WAF rule:
{
"description": "Block observed 7Siters crawler",
"expression": "lower(http.user_agent) contains \"7siters\"",
"action": "block"
}
For Nginx, scope the block to the whole site only when that is the intended decision; otherwise protect private and high-cost routes first:
map $http_user_agent $block_7siters {
default 0;
~*7Siters 1;
}
server {
location ~ ^/(private|internal|account|uploads|api)/ {
if ($block_7siters) { return 403; }
try_files $uri $uri/ =404;
}
}
A full-site block can be expressed at the server boundary, but test it carefully:
if ($block_7siters) { return 403; }
A User-Agent rule is easy to spoof or evade. Do not create an IP allowlist or denylist without source evidence; the reviewed 7Siters page publishes no ranges. Test browsers, social previews, feed readers, search crawlers, and approved monitors. Pair edge matching with authentication, rate limits, signed assets, and the documented 7Siters.txt removal workflow.
Review checklist
Search logs for 7Siters and preserve the full header, source IP, paths, response sizes, statuses, timing, and request rate. Decide whether you want to opt out of published statistics, stop crawling entirely, or only protect private and high-cost routes.
If you choose the first-party workflow, create the root-level 7Siters.txt file with exactly delete, submit the domain in the site's removal form without a protocol, and retain evidence of the request. For immediate enforcement, use the documented Apache rule or a narrow WAF/Nginx match. Publish a User-agent: 7Siters robots group as a communication layer, not as a security boundary.
Re-test public pages, internal links, feeds, sitemaps, media, uploads, and APIs separately. Monitor for continued requests after the opt-out and contact support@7siters.com with timestamps, source IPs, full User-Agent, domain, and representative logs if the documented removal process does not work.
References
- 7Siters — 7ooo.ru — first-party page describing website statistics, collected data, the
7Siters.txtremoval file, Apache block example, and support address. - 7Siters domain-removal form — the same first-party page instructs owners to enter the domain without
http://orhttps://and use a root-level7Siters.txtcontainingdelete. - Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.