SBIntuitionsBot: Robots.txt & Crawl Policy Reference
Technical reference for SBIntuitionsBot, SB Intuitions' crawler for AI development and information analysis. Learn its robots policy and distinguish Sarashina search fetches.
AI Summary: SB Intuitions officially documents
SBIntuitionsBotas a crawler used for AI development and information analysis, and states that its crawlers comply with RFC 9309's Robots Exclusion Protocol. It separately documentsSBIntuitions-SearchBotfor user questions in the Sarashina service. Use the exact bot group, distinguish the two purposes, and verify observed requests from logs because no public IP range is provided.
Role and policy boundary
SB Intuitions Corp.'s official crawler page lists SBIntuitionsBot for AI development and information analysis. This is a first-party purpose statement and is materially different from the company's separately listed SBIntuitions-SearchBot, which may fetch pages on behalf of a user asking Sarashina a question and use the retrieved information for inference. The search bot's referenced information is stated not to be used for AI development.
The registry label includes a browser-compatible suffix, but the official page identifies the stable token as SBIntuitionsBot. Match that token rather than requiring the entire versioned header. Blocking it may reduce collection for SB Intuitions' AI-development workflows; it does not control the separate Sarashina search fetcher, general search engines, or unrelated AI crawlers.
SB Intuitions states that crawlers it operates comply with RFC 9309. To exclude all paths from the AI-development crawler:
User-agent: SBIntuitionsBot
Disallow: /
To exclude selected paths while leaving public documentation available:
User-agent: SBIntuitionsBot
Allow: /docs/
Allow: /public-reference/
Disallow: /private/
Disallow: /internal/
Disallow: /uploads/
The official page says that policy changes may take time to propagate. Robots.txt remains advisory and is not an access-control mechanism; private content needs authentication and authorization.
Layered verification
Start with access logs and search for the stable SBIntuitionsBot token. Preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. SB Intuitions does not publish a source IP list on the reviewed page, so do not invent an allowlist or treat a header as cryptographic identity.
Keep SBIntuitionsBot and SBIntuitions-SearchBot separate in telemetry and policy. Their documented purposes differ: one supports AI development and information analysis, while the other responds to user questions and is not used for AI development according to the official page. A request pattern alone cannot prove which downstream use occurred, but the exact token provides a useful operational classification.
Evaluate /robots.txt independently. Confirm the canonical host, status, content type, exact group, path match, and propagation time. Page-level directives are separate signals:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives do not protect private content and should not be confused with a model-training opt-out when the operator's published policy is specific to its crawler token.
WAF and Nginx remediation examples
If you want to block the AI-development crawler, use the stable token in a narrow WAF rule:
{
"description": "Block SB Intuitions AI-development crawler",
"expression": "lower(http.user_agent) contains \"sbintuitionsbot\"",
"action": "block"
}
Do not reuse that rule for SBIntuitions-SearchBot; decide its user-requested search policy separately. For Nginx, scope the block to high-cost or sensitive routes while validating the rollout:
map $http_user_agent $block_sbintuitionsbot {
default 0;
~*SBIntuitionsBot 1;
}
server {
location ~ ^/(private|internal|uploads|originals|api)/ {
if ($block_sbintuitionsbot) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent rule can be spoofed and may miss a client using another header. Test ordinary browsers, social previews, approved monitors, Sarashina search requests, and the separate search token. Combine edge rules with authentication, rate limits, and origin protection.
Review checklist
Decide whether AI-development collection is allowed for your public content and whether the separate Sarashina search experience should remain available. Publish an exact User-agent: SBIntuitionsBot group, allow or disallow only the intended paths, and record that the operator says propagation may take time.
Search logs for both official tokens and keep their records separate. Confirm status, response size, path, source network, and timing; do not create IP allowlists because the reviewed policy publishes none. Protect private pages, uploads, APIs, and customer data with authentication regardless of robots preferences. Re-test public docs and excluded paths after the policy has had time to propagate.
References
- SB Intuitions crawler policy — official Japanese page listing
SBIntuitionsBot,SBIntuitions-SearchBot, purposes, RFC 9309 compliance, and robots example. - RFC 9309 — standard for the Robots Exclusion Protocol referenced by SB Intuitions.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.