Bot directory / search-engine

Y!J: Robots.txt & Crawl Policy Reference

Technical reference for the Yahoo! Japan search engine crawler, noting its documented purpose but lack of a published IP verification list.

AI Summary: Y!J refers to the web crawlers operated by Yahoo! Japan for its search engine and related services. Official documentation confirms that its crawlers (such as Y!J-ASR and Y!J-BRW) respect robots.txt directives. However, because Yahoo! Japan does not publish a strict static IP list or a specific reverse-DNS verification method for these crawlers, use standard network validation and behavioral analysis before trusting the header.

Role and policy boundary

Yahoo! Japan operates web crawlers to index content for its search engine and related services. The registry records the historical User-Agent as Y!J-ASR/0.1 crawler. The official documentation confirms that Yahoo! Japan crawlers respect robots.txt directives targeted at tokens such as Y!J-ASR or Y!J-BRW.

Because the operator does not publish a verifiable IP range or reverse-DNS procedure, this profile is marked documented-limit. The User-Agent string is easily spoofed, and an undocumented IP cannot be cryptographically authenticated. Do not assume that every request bearing this token is legitimate or that the token grants permission for AI input, model training, or data retention beyond general search indexing.

To exclude this crawler across your entire site, add the following to your robots.txt:

configuration / code
User-agent: Y!J-ASR
User-agent: Y!J-BRW
Disallow: /

For selective access:

configuration / code
User-agent: Y!J-ASR
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/

These are explanatory site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. The official documentation does not publish a Yahoo! Japan crawler IP range, reverse-DNS procedure, Crawl-delay, rate guidance, or verification workflow. The User-Agent is useful for triage but is not authentication.

Do not treat the word Y!J in a header as proof of ownership. Check source IP and DNS evidence independently, and require behavioral consistency before creating a trust exception.

Compare observed behavior with a search-indexing hypothesis without turning it into attribution. Public HTML, metadata, feeds, sitemaps, and ordinary assets may be requested by many tools. Private endpoints, licensed content, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not Yahoo! Japan identity or downstream use.

Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish. An absent IP list is not evidence of robots non-compliance, but it prevents proactive allowlisting.

Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate an undocumented crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes search indexing, AI input, reference use, and model training, document each purpose separately rather than inferring permission from a registry label.

WAF and Nginx remediation examples

Because the identity cannot be cryptographically verified via published IP lists, use report-only logging and a narrow observation. Avoid broad yahoo or browser rules that can affect unrelated clients:

configuration / code
{
  "description": "Observe unverified Yahoo Japan candidates",
  "expression": "lower(http.user_agent) contains \"y!j\"",
  "action": "log"
}

After independent verification and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule:

configuration / code
map $http_user_agent $block_yj_private {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*Y!J-ASR 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_yj_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof. Do not invent an IP allowlist, Yahoo ASN exception, reverse-DNS suffix, rate, training policy, or permanent trust exception. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact Y!J substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. The operator does not provide a verification method.

Re-check the crawler registry, the Yahoo! Japan operator site, robots.txt, and operator contact path. Keep this profile at documented-limit until a current source publishes a source-verification method, rate guidance, or opt-out process.

Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.

Review search indexing, AI input, reference use, and model-training decisions separately. The documentation does not establish downstream permission or use. Do not claim successful blocking or verification from configuration alone; validate later logs.

Record the date and reason for the documented-limit classification so future evidence can be compared. If a new deployment identifies itself differently, create a separate evidence trail.

References

  1. Yahoo! Japan Help: Crawler User Agents — official documentation on Yahoo! Japan's crawler user agents.
  2. Crawler User Agents community registry — registry context for the Y!J label.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.