ichiro: Robots.txt & Crawl Policy Reference
Technical reference for the historical ichiro/mobile goo crawler label, with explicit limits around its unavailable documentation and current verification.
AI Summary:
ichiro/mobile goois a historical registry label with a legacy mobile-style User-Agent. The linked goo help article failed DNS resolution over both HTTP and HTTPS during review, so no current operator policy, robots behavior, IP range, or verification method was established. Treat matching traffic as an unverified historical observation.
Role and policy boundary
The inventory associates ichiro with the goo search engine and records this User-Agent:
DoCoMo/2.0 P900i(c100;TB;W24H11) (compatible; ichiro/mobile goo; +http://help.goo.ne.jp/help/article/1142/)
The linked help article could not be reached in the headless browser: both http://help.goo.ne.jp/help/article/1142 and its HTTPS equivalent returned net::ERR_NAME_NOT_RESOLVED. That prevents this review from confirming whether the historical client is still operated, whether goo changed its identity, or whether current robots and verification instructions exist elsewhere.
The registry’s search-crawler description is a historical role hypothesis only. A request bearing the header may come from an old deployment, a test client, a fork, or a spoofed request. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private content from the label.
If your logs confirm this exact token and your policy is to exclude it from discovery, a narrow defensive group could be:
User-agent: ichiro/mobile goo
Disallow: /
For selective access to approved public pages:
User-agent: ichiro/mobile goo
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
These are site-owner examples, not recovered goo instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No current goo or ichiro source reviewed for this profile published an IP range, reverse-DNS method, rate policy, or current canonical header. Do not treat the historical mobile header alone as authentication.
Compare observed behavior with a search-indexing hypothesis without treating it as attribution. Public HTML, canonical metadata, feeds, sitemaps, and ordinary assets may be consistent with indexing. Private endpoint access, high concurrency, repeated retries, unexpected downloads, or traffic that ignores your site policy may indicate spoofing, abuse, a fork, or an unrelated client. These observations establish impact, not operator identity or downstream use.
Evaluate /robots.txt independently if you publish a rule. Confirm that the response is served by the intended host, returns a successful status and text content type, and contains the exact group you intend to apply. Test the complete registered header and the group token separately, because robots matching is not a substitute for request authentication. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate a historical goo client or secure private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If logs show a repeatable unwanted token, start in report-only mode and preserve representative requests. Adapt the expression to your WAF provider; it identifies a historical header pattern only:
{
"description": "Review historical ichiro/mobile goo traffic",
"expression": "lower(http.user_agent) contains \"ichiro/mobile goo\"",
"action": "log"
}
After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:
map $http_user_agent $review_ichiro {
default 0;
~*ichiro/mobile[ ]goo 1;
}
server {
location ~ ^/(admin|account|private|internal|api)/ {
if ($review_ichiro) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not invent an IP allowlist, reverse-DNS suffix, or current goo exception. A User-Agent is easy to spoof and the mobile-style match could catch an authorized compatibility client. Test public HTML, feeds, sitemaps, media, uploads, account flows, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for the complete registered header, including the old help URL. Record representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the source can be independently verified and whether observed behavior actually resembles public search indexing.
Re-check the goo help domain and the registry for a current operator source when the host becomes reachable. Keep this profile at legacy-label until a credible source publishes a current policy, canonical User-Agent, source-verification method, or robots guidance. Treat the DNS failure as a source-access limitation, not proof that the crawler is inactive.
Decide whether your objective is to preserve search visibility, limit extraction, protect private material, or reduce crawl load. Publish exact robots rules and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent access logs and response behavior.
References
- goo help article for ichiro — registry-linked source; direct headless-browser request failed with
net::ERR_NAME_NOT_RESOLVEDon 2026-08-25. - goo help article HTTPS equivalent — direct headless-browser request also failed with
net::ERR_NAME_NOT_RESOLVEDon 2026-08-25. - Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.