Bot directory / ai-training

cohere-ai: Robots.txt & Crawl Policy Reference

Technical reference for the registry-listed cohere-ai label. Learn what Cohere currently documents about web crawlers and how to avoid confusing product APIs with a public crawler.

AI Summary: cohere-ai is a registry label, not a currently documented Cohere training User-Agent. Cohere’s official crawler page states that it does not use Cohere bots or User-Agents to crawl or scrape web content for training generative AI foundation models at the time of publication, lists no active training bot, and requires future crawlers to respect robots.txt. Treat any observed cohere-ai request as unverified until its exact token and source are confirmed.

Role and policy boundary

Cohere’s public documentation separates its AI products from its crawler policy. Cohere builds models and products such as Command, Embed, Rerank, and North, but those product descriptions do not establish that a public web crawler called cohere-ai is currently operating. A customer can call Cohere APIs, connect private data to an application, or run a separate ingestion workflow without sending a request to a website as a Cohere crawler.

The official Cohere Web Crawlers page makes a more specific statement: at the time of the document, Cohere does not use Cohere bots or User-Agents to crawl or scrape web content to train generative AI foundation models. Its table lists no active training bot or User-Agent. Cohere also says that its policy requires crawlers to be designed to respect robots.txt, while showing Coherebot as a generic future-blocking example. That example is not evidence that cohere-ai or Coherebot is an active crawler today.

For a site owner, the first policy is therefore not to invent an allow or block rule around an unverified token. Inspect access logs, identify the complete User-Agent, check the source address, and determine whether the request is actually a Cohere crawler, an unrelated client, or a product integration using a different identity. Keep private content behind authentication regardless of the robots result.

If the exact cohere-ai token is observed and your policy is to block that self-declared identifier, you may express the preference narrowly:

configuration / code
User-agent: cohere-ai
Disallow: /

If you later confirm an active Cohere crawler from updated vendor documentation, use the exact User-Agent named by Cohere rather than copying the registry label. Do not block the entire Cohere product surface or all API clients based on a website crawler assumption.

A robots rule communicates a crawl preference; it does not authenticate a caller or prevent an authenticated application from receiving data. It also does not prove that a request is being used for foundation-model training.

Layered verification

Start with the complete request record: User-Agent, source IP, method, path, status, redirect chain, request rate, and response size. The current Cohere crawler documentation does not publish a definitive cohere-ai User-Agent, so the registry slug cannot serve as vendor authentication. If the observed request has no stable token or source evidence, classify it as unknown rather than assigning it to Cohere.

Next, inspect the canonical top-level /robots.txt and confirm whether a dedicated group matches the exact observed token. A wildcard rule may apply to the request, but it should not be interpreted as proof of operator identity. Compare the result with the current Cohere documentation and record the retrieval date because the vendor may publish a new crawler table later.

The Policy Engine should preserve this evidence boundary. A match to cohere-ai can explain a site-specific rule, but the absence of a public Cohere specification should remain visible in the report. Do not convert a product capability such as RAG, semantic search, or an agent platform into a claim that Cohere is crawling the public web for training.

Page-level directives are separate signals:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives communicate indexing preferences but do not prevent access or provide authentication. Use application authorization, WAF rules, and rate limiting when the concern is confidentiality or resource consumption.

WAF and Nginx remediation examples

If your logs show the exact unverified token and your policy is to block it, a narrow WAF expression can reduce repeat traffic. Label the decision as a self-declared-token match, not proof that the request belongs to Cohere:

configuration / code
{
  "description": "Block declared cohere-ai registry token",
  "expression": "lower(http.user_agent) contains \"cohere-ai\"",
  "action": "block"
}

A path-specific Nginx example is:

configuration / code
map $http_user_agent $deny_cohere_ai {
    default 0;
    ~*cohere-ai 1;
}

server {
    location ~ ^/(internal|customer-data|account)/ {
        if ($deny_cohere_ai) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Do not use this match as the only protection for confidential data, and do not treat Cohere’s generic Coherebot example as an instruction to block a token that your logs do not show. If a future official table names a crawler for model training, update the rule to the exact documented token and re-evaluate the business purpose separately from product API traffic.

Review checklist

Request /robots.txt from each canonical subdomain and verify the status, content type, final URL, and exact rule applied to the observed token. Test one public URL, one private URL, and one redirecting URL. Record the complete User-Agent, source address, HTTP status, redirects, response headers, response size, and request rate.

Then re-read Cohere’s current crawler page and confirm whether its table still lists no active training bot. Keep cohere-ai classified as unverified unless a primary source names it. Re-test after CDN, WAF, origin, or robots changes, and never claim that Cohere is training on your public content from a token that the vendor has not documented.

References

  1. Cohere Web Crawlers — official crawler policy, current training-crawler statement, and robots.txt guidance.
  2. Welcome to Cohere — official overview of Cohere models and products, including RAG, semantic search, and agents.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.