← Bot Directory/Mail.RU_Bot
Bot directory / search-engine

Mail.RU_Bot: Robots.txt & Crawl Policy Reference

Technical reference for the Mail.RU_Bot registry label, including the observed go.mail.ru robots policy and limits around current bot-specific verification.

AI Summary: Mail.RU_Bot/2.0 is a community-registry label associated with Mail.RU search crawling. The historical help URL redirects to the current Mail help service, while go.mail.ru/robots.txt exposes a global allowlist for selected paths followed by Disallow: /; no Mail.RU_Bot-specific group, IP range, Crawl-delay, or DNS verification method was published. Treat the registered header as a signal, not authentication.

Role and policy boundary

The registry describes Mail.RU_Bot as a Mail.RU search-engine web crawler and records this User-Agent:

configuration / code
Mozilla/5.0 (compatible; Linux x86_64; Mail.RU_Bot/2.0; +http://go.mail.ru/help/robots)

The historical documentation path redirected to the current https://help.mail.ru/ service-help homepage. The direct https://go.mail.ru/robots.txt endpoint is live and contains a global User-agent: * group. It allows the root and selected paths including /search_images, /images, /search_video, /video, /news, /help, /help.html, /addurl, /blog, and certain query patterns, then includes Disallow: /.

That file does not publish an exact Mail.RU_Bot group, a Crawl-delay, an IP range, a reverse-DNS verification method, or an opt-out contact in the reviewed content. Do not infer that the global rules authenticate or fully describe the historical Mail.RU_Bot client. Search indexing is the registry’s role hypothesis, not a current first-party confirmation of the specific header.

A matching request may come from a legacy client, a Mail/RU partner, a test harness, or a spoofed header. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private mail and cloud data from the token.

If logs confirm this exact token and your policy is to exclude it from discovery, a narrow defensive group could be:

configuration / code
User-agent: Mail.RU_Bot
Disallow: /

For selective access to public pages while protecting account and service endpoints:

configuration / code
User-agent: Mail.RU_Bot
Allow: /public/
Allow: /docs/
Disallow: /account/
Disallow: /mail/
Disallow: /cloud/
Disallow: /private/
Disallow: /api/

These are site-owner examples, not recovered Mail.RU instructions. Robots.txt is advisory and cannot protect private mail, cloud files, or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. The reviewed Mail/RU endpoint publishes no Mail.RU_Bot-specific IP range, reverse-DNS procedure, rate policy, or current canonical bot page. The header is useful for triage but is not authentication.

Compare observed behavior with a search-indexing hypothesis without turning it into attribution. Public HTML, search-visible documents, feeds, sitemaps, images, and video metadata may be consistent with a search service. Access to mail, cloud files, accounts, private APIs, high concurrency, repeated retries, unexpected downloads, or traffic that ignores your restrictions may indicate spoofing, abuse, a partner integration, or another client. These observations establish impact and data risk, not operator identity or downstream use.

Evaluate /robots.txt independently. Confirm that it is served from the intended host, returns a successful status and text content type, and contains the exact group you intentionally publish. The current global rule allows selected Mail search/help paths but then disallows /; do not generalize this observed file to other Mail/RU-owned domains or to a different historical endpoint. Test the full header, the token, and the global group separately. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate Mail.RU_Bot or secure private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

If logs show the registry token, begin with a narrow report-only rule and correlate it with source evidence. Adapt the expression to your WAF provider; it identifies an observed header only:

configuration / code
{
  "description": "Observe Mail.RU_Bot search candidates",
  "expression": "lower(http.user_agent) contains \"mail.ru_bot\"",
  "action": "log"
}

After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:

configuration / code
map $http_user_agent $block_mail_ru_private {
    default 0;
    ~*Mail\.RU_Bot/2\.0 1;
}

server {
    location ~ ^/(account|mail|cloud|private|internal|api)/ {
        if ($block_mail_ru_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and may catch a legitimate Mail/RU integration or a local test client. Do not invent an IP allowlist, reverse-DNS suffix, or rate exception from the global robots file. Test public pages, images, video metadata, feeds, sitemaps, mail, cloud, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, data-loss monitoring, and anomaly detection.

Review checklist

Search logs for the complete registered header and the simpler Mail.RU_Bot/2.0 token. Preserve representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the source can be independently verified and whether requested content is public; do not classify it as authentic from the header alone.

Re-check the historical help path, current Mail/RU crawler documentation, and go.mail.ru/robots.txt for an exact Mail.RU_Bot group before relying on any global rule. Confirm that accounts, mail, cloud data, private content, and APIs are protected by application authorization, not robots.txt.

Keep this profile at documented-limit until a first-party page confirms the registered header, current search role, source-verification method, IP ranges, rate guidance, or opt-out process. Treat the redirect as an endpoint migration observation, not proof that the historical client is still active.

If you need to change search visibility, publish an exact robots group only after confirming the observed token and verify subsequent access logs. Do not claim that the global Disallow: / rule controls Mail.RU_Bot specifically without matching semantics and current client evidence.

References

  1. go.mail.ru robots.txt — official current robots file observed with the headless browser on 2026-08-25; it exposes a global selected-path allowlist followed by Disallow: / and no Mail.RU_Bot-specific group in the reviewed content.
  2. Historical Mail.RU robots help URL — the registry-linked http://go.mail.ru/help/robots redirected to the current Mail help homepage during review; no bot-specific policy was inferred.
  3. Crawler User Agents community registry — community source for the registry header; it does not authenticate current Mail/RU infrastructure.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.