IbouBot: Robots.txt & Crawl Policy Reference
Technical reference for IbouBot, including its published search-crawler policy, Crawl-delay guidance, forward-confirmed reverse DNS verification, and current IP range.
AI Summary: Ibou documents IbouBot as the crawler for its web-graph search engine. It says the bot follows robots.txt and Crawl-delay, uses a 5-second same-host politeness delay and 2.5-second same-IP delay, and publishes
217.113.196.0/24with forward-confirmed reverse-DNS verification underibou.io. The published User-Agent is useful for triage but is not authentication by itself.
Role and policy boundary
Ibou’s official IbouBot page identifies the client as a crawler that discovers publicly accessible pages and indexes them for Ibou’s search engine. The page says the purpose is to help readers find pages and return visitors to publishers. It also states that Ibou does not train AI models with the crawled data. That is a statement about the documented Ibou service and should not be generalized to other clients using the header.
The published User-Agent is:
Mozilla/5.0 (compatible; IbouBot/1.0; +bot@ibou.io; +https://ibou.io/iboubot.html)
Ibou explicitly says that it follows robots.txt and its disallow directives. It documents a 5-second delay between two queries on the same host and a 2.5-second delay between two queries on the same IP of the same domain, with a worst-case load of approximately 0.4 request/second per domain, or about 24 requests per minute. Site owners can request a longer interval or a stop through the published contact address.
To prevent IbouBot from crawling the WordPress admin area, Ibou gives this form:
User-agent: IbouBot
Disallow: /wp-admin/
To stop the documented crawler from accessing the whole site:
User-agent: IbouBot
Disallow: /
For a selective public-search boundary:
User-agent: IbouBot
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
Crawl-delay: 10
Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries. A site owner may also contact bot@ibou.io with the domain, source IPs, and representative log lines to request a slower crawl or an exclusion.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Ibou documents forward-confirmed reverse DNS as the reliable way to distinguish genuine traffic from a spoofed header.
For an observed source IP, perform a reverse lookup and confirm that the returned hostname ends with ibou.io. Ibou’s example maps 217.113.196.1 to c001.ibou.io. Then perform a forward lookup of that hostname and confirm that it resolves back to the original 217.113.196.1. A mismatch means the request should not be treated as genuine IbouBot traffic.
The official IP JSON currently reports 217.113.196.0/24 and includes a creation time of 2025-07-25T14:16:44.000000. Ibou warns that the list may change, so fetch it periodically rather than copying it as a permanent allowlist or blocklist. Use the current JSON plus DNS checks and the full User-Agent; never rely on only one signal.
Compare request behavior with the documented search role. Publicly accessible pages and moderate, rate-limited traversal may be consistent with indexing. Private endpoint access, high concurrency, repeated retries, unexpected downloads, or traffic from outside the documented range with failed DNS confirmation may indicate spoofing, a changed deployment, or another client. These observations establish impact, not downstream use.
Evaluate /robots.txt independently. Confirm that it is served from the correct host, returns a successful status and text content type, contains an exact IbouBot group, and applies Crawl-delay and path rules as intended. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not secure private routes or authenticate the client. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If you want to recognize Ibou without maintaining a stale IP list, begin with logging and DNS validation. Adapt the expression to your WAF provider; do not let the User-Agent alone authorize access:
{
"description": "Observe IbouBot after DNS and range verification",
"expression": "lower(http.user_agent) contains \"iboubot\"",
"action": "log"
}
If the policy is to restrict verified IbouBot traffic on sensitive routes, use a narrow match and combine it with your provider’s verified-bot signal where available:
map $http_user_agent $block_iboubot_private {
default 0;
~*IbouBot/1\.0 1;
}
server {
location ~ ^/(admin|account|private|internal|api)/ {
if ($block_iboubot_private) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not use the Nginx header match as proof of origin. Validate the source IP using reverse and forward DNS, refresh the official JSON range, and review false positives before enforcement. Test public pages, feeds, sitemaps, media, uploads, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for the complete published header and record source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, and redirect chain. Confirm that the source falls within the current 217.113.196.0/24 range and that DNS is forward-confirmed before classifying it as genuine IbouBot.
Check that your exact IbouBot robots group and Crawl-delay express the intended search-visibility trade-off. Protect private routes with application authorization, not robots.txt. If the bot is too frequent, apply a documented delay or contact bot@ibou.io with representative evidence rather than relying on a broad emergency block.
Refresh the official IP JSON periodically and re-check the IbouBot page for changes to the User-Agent, rate policy, Cloudflare Verified Bot status, or contact process. Do not claim successful blocking from a configuration change alone; verify subsequent access logs and response behavior.
Remember that Ibou’s statement that it does not train AI models applies to the documented Ibou service, not to arbitrary clients impersonating its header. Keep attribution tied to the verification evidence.
References
- IbouBot crawler documentation — official page documenting purpose, User-Agent, robots.txt/Crawl-delay behavior, politeness limits, DNS verification, Cloudflare classification, contact process, and search-versus-training statement; reviewed with the headless browser on 2026-08-25.
- IbouBot IP ranges JSON — official current range document, reporting
217.113.196.0/24and its creation time; reviewed with the headless browser on 2026-08-25. - Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.