Bot directory / ai-training

AI2Bot: Robots.txt & Crawl Policy Reference

Guide to AI2Bot, the web crawler operated by the Allen Institute for AI. Understand its role in open language model training and how to configure access.

AI Summary: AI2Bot (and variants like Ai2Bot-Dolma) is a web crawler operated by the Allen Institute for AI (AI2), a non-profit research institute. It collects publicly available web content to build open datasets used for training open-source language models. It officially respects robots.txt directives.

Role and policy boundary

AI2Bot scrapes the web strictly for AI research and the development of open-source foundational models (such as the Dolma dataset). Unlike commercial crawlers (e.g., GPTBot or ClaudeBot), the data collected by AI2Bot contributes to open science and the broader AI research community.

However, if your organization's policy is to opt out of all AI training—regardless of the crawling entity's non-profit status or open-source contributions—you must explicitly block this bot.

A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting. Start with a dedicated group:

configuration / code
User-agent: AI2Bot
Allow: /
Disallow: /staging/
Disallow: /internal/

To stop access for the entire site, use:

configuration / code
User-agent: AI2Bot
Disallow: /

Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect assistant and training crawlers too, and it makes later audits harder because the source of the decision is less specific.

Layered verification

Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.

The Allen Institute for AI explicitly states that AI2Bot respects robots.txt directives. You can opt out of their data collection by adding a Disallow rule specifically for the AI2Bot User-Agent.

The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why a bot is allowed while another is blocked, rather than returning one blended website score.

Page-level directives can still override the intended outcome for indexing:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

If a response uses these tags, the report marks the result as blocked or conflicting even if the crawler-specific robots group is permissive. This is especially important for canonical pages served through an edge cache where headers may differ from the origin response.

WAF and Nginx remediation examples

If you prefer to manage bot traffic at the network edge to conserve server resources, or to implement a blanket ban on all AI training crawlers, you can block AI2Bot using a standard WAF rule:

configuration / code
{
  "description": "Block Allen Institute AI2Bot",
  "expression": "lower(http.user_agent) contains \"ai2bot\"",
  "action": "block"
}

Use your platform's actual middleware response pattern rather than copying this simplified example without review. Never place a secret, verification token, or internal policy identifier in a public header.

configuration / code
map $http_user_agent $block_ai2bot_private {
    default 0;
    ~*AI2Bot 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_ai2bot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Review checklist

Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.

Verify that the dedicated group appears before relying on a wildcard, test a representative public and private path, and compare live response headers with robots.txt. Keep the policy close to the content owner's intent and record whether the site wants discovery, citation, or no access at all.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.