← Bot Directory/cohere-training-data-crawler
Bot directory / ai-training

cohere-training-data-crawler: Robots.txt & Crawl Policy Reference

Technical reference for the registry-listed Cohere training crawler label. Compare the observed token with Cohere's current official statement before applying an opt-out.

AI Summary: cohere-training-data-crawler (+crawler@cohere.ai) is a registry-listed label for an alleged Cohere training crawler, but Cohere’s current official crawler page says it does not use Cohere bots or User-Agents to crawl or scrape web content for training generative AI foundation models at the time of publication and lists no active training bot. Treat the token as legacy or unverified, verify it in logs, and do not attribute it to current Cohere activity without fresh primary evidence.

Role and policy boundary

The registry label describes a crawler that would collect public web content for AI model development. That is a meaningful policy category because a site owner may want public pages to remain searchable while opting out of foundation-model training. However, the current Cohere source reviewed for this profile directly limits that interpretation: Cohere says it does not use Cohere bots or User-Agents for that training purpose at the time of publication, and its active-bot table lists N/A.

This creates an evidence conflict that should be visible in an audit. A User-Agent containing cohere.ai can be copied by unrelated software, a legacy crawler can disappear while a registry entry remains, and an external directory can preserve an old label after a vendor changes policy. The correct first step is to inspect the exact request and compare it with Cohere’s current documentation, not to assume that a training pipeline is active.

If the token is genuinely present in your logs and your organization wants to block that self-declared identifier, a narrow robots rule can express the preference:

configuration / code
User-agent: cohere-training-data-crawler
Disallow: /

If a future Cohere page names a different active training crawler, use the exact vendor-published token and keep that rule separate from search or user-directed retrieval. Cohere’s generic Coherebot example should not be treated as proof that this registry token is current.

A robots rule is a crawl preference, not authentication. It does not prevent an authenticated application from receiving data, protect a private route, or establish the downstream purpose of a request. Keep customer, account, staging, and transactional content behind application authorization.

Layered verification

Start with the complete HTTP evidence: User-Agent, source IP, method, path, timestamp, status, redirect chain, response size, and request rate. Compare the source address against any current vendor information and investigate whether crawler@cohere.ai is merely part of a self-declared string. The Cohere crawler page currently does not confirm cohere-training-data-crawler, so the token must remain an observation rather than a verified Cohere identity.

Then fetch the canonical top-level /robots.txt and test the exact path scope. Check that the policy file is publicly reachable, served with the expected content type, and not replaced by a login or challenge response. Compare the result with any page-level metadata and response headers, but keep each signal independent. A robots deny result records the site owner’s preference; it does not prove that a future request will stop or that the caller is actually Cohere.

The Policy Engine should report the conflict explicitly: the registry classifies this label as training-related, while the current official Cohere page says no Cohere training bot is in use at the time of its publication. If the request is not observed, do not add a block rule solely because a third-party inventory lists the token. If it is observed, preserve the raw evidence and re-check the vendor source before escalating the decision.

Page-level directives are separate signals:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives address indexing preferences, not authentication or model-training provenance. Use access controls, WAF rules, and rate limiting for confidentiality and resource protection.

WAF and Nginx remediation examples

If your logs show the exact token and the policy is to block it, a narrow WAF expression can reduce repeat requests. Label it as a self-declared-token match and avoid treating it as a proof of Cohere ownership:

configuration / code
{
  "description": "Block declared Cohere training-crawler token",
  "expression": "lower(http.user_agent) contains \"cohere-training-data-crawler\"",
  "action": "block"
}

A path-specific Nginx example is:

configuration / code
map $http_user_agent $deny_cohere_training_crawler {
    default 0;
    ~*cohere-training-data-crawler 1;
}

server {
    location ~ ^/(internal|customer-data|account)/ {
        if ($deny_cohere_training_crawler) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Do not use this match as the only protection for private data. Do not block all Cohere API traffic or all traffic containing cohere.ai based on this registry label. If Cohere publishes an active crawler later, update the rule to the exact documented token and document the new source date.

Review checklist

Request /robots.txt from each canonical subdomain and verify the status, content type, final URL, and exact rule applied to the observed token. Test one public URL, one private URL, and one redirecting URL. Record the User-Agent, source address, HTTP status, redirects, response headers, response size, and request rate.

Re-read Cohere’s current crawler policy and confirm whether its active-bot table has changed. If the request is observed, preserve the raw header and network evidence; if it is not observed, do not treat the registry entry as proof of current collection. Re-test after CDN, WAF, origin, or robots changes, and never claim that Cohere is training on a site’s content from this unconfirmed token alone.

References

  1. Cohere Web Crawlers — official statement about current crawler use, robots.txt policy, and the absence of a listed active training bot.
  2. Welcome to Cohere — official overview of Cohere models and products, distinct from crawler identity.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.