coccocbot: Robots.txt & Crawl Policy Reference
Technical reference for coccocbot, the official web crawler for the Cốc Cốc search engine. Learn how to verify its traffic and manage its access.
AI Summary:
coccocbotis the official web crawler for Cốc Cốc, a major search engine in Vietnam. It crawls the web to build the Cốc Cốc search index, operating through several specialized variants (such ascoccocbot-webandcoccocbot-image). Because it is a major regional search engine, blocking coccocbot will prevent your site from appearing in Cốc Cốc search results. Verify its traffic using reverse DNS lookups to ensure the IPs belong to.coccoc.com.
Role and policy boundary
According to official documentation, Cốc Cốc operates several specialized crawlers, all generally falling under the coccocbot umbrella. The primary web crawler identifies itself as coccocbot-web, while other variants handle images (coccocbot-image), fast crawling (coccocbot-fast), ads (coccocbot-ads), and shopping data (coccocbot-shopping).
As a major search engine crawler, its primary purpose is public indexing to serve users in Vietnam. Blocking coccocbot means your content will not be discoverable by users searching on Cốc Cốc. Do not infer that its primary purpose is AI training, though like all major search engines, indexed data may inform their ecosystem.
If your logs confirm an exact coccocbot token and you want to prevent your site from appearing in Cốc Cốc search results, publish:
User-agent: coccocbot
Disallow: /
For selective access (e.g., hiding private sections from search):
User-agent: coccocbot
Allow: /
Disallow: /private/
Disallow: /internal/
Disallow: /api/
Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Because coccocbot is a search engine crawler, it may be spoofed by malicious scrapers, so you must verify its authenticity.
Cốc Cốc provides a clear verification method via reverse DNS. Perform a reverse DNS (PTR) lookup on the accessing IP address. A genuine coccocbot IP will resolve to a hostname ending in .coccoc.com. You can then perform a forward DNS lookup on that hostname to ensure it matches the original IP.
Analyze behavior without assigning purpose prematurely. Requests for public pages, feeds, sitemaps, and metadata resemble authorized search indexing; deep traversal of private areas, high concurrency ignoring crawl delays, repeated retries, or private API access may indicate spoofing. These patterns demonstrate operational impact but cannot prove the operator or downstream use.
Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:
<meta name="coccocbot" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.
WAF and Nginx remediation examples
Once logs confirm an exact unwanted token (e.g., a spoofed bot failing reverse DNS verification, or if you intentionally want to block Cốc Cốc globally), a narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:
{
"description": "Block observed coccocbot token",
"expression": "lower(http.user_agent) contains \"coccocbot\"",
"action": "block"
}
For Nginx, scope enforcement to private and high-cost routes while investigating public access:
map $http_user_agent $block_coccocbot {
default 0;
~*coccocbot 1;
}
server {
location ~ ^/(private|internal|account|uploads|paywall|api)/ {
if ($block_coccocbot) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent rule is easy to spoof or evade and may block a legitimate client using the substring. Do not create an IP allowlist or broad network block without current operator evidence. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.
Review checklist
Search logs for every exact header that may be associated with coccocbot (including -web, -image, -fast, etc.) and preserve representative requests. Record paths, response sizes, statuses, source networks, timing, and rate. Verify the traffic by performing a reverse DNS lookup to ensure the hostname ends in .coccoc.com.
Decide whether your objective is to preserve search visibility in Vietnam, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group only for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately.
References
- Cốc Cốc Search Console: Các robot của Cốc Cốc — official documentation portal detailing the crawler variants and reverse DNS verification, accessed 2026-08-24.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.