psbot: Robots.txt & Crawl Policy Reference
Technical reference for the historical PicSearch psbot-image registry label, documenting the discontinued service and limits around current crawler verification.
AI Summary: The registry identifies
psbot-imageas a PicSearch image-discovery crawler, but its linked bot page returns 404 and the current PicSearch homepage states that the service ended in 2022. No current psbot-image policy, robots rule, rate limit, IP range, reverse-DNS method, or opt-out path is verified. Treat the token as a historical signal and investigate any present-day match as an unattributed client.
Role and policy boundary
The registry describes psbot as a PicSearch web crawler for image discovery and records this User-Agent:
psbot-image (+http://www.picsearch.com/bot.html)
The linked bot page returned 404 Not Found during headless-browser review. The current https://www.picsearch.com/ homepage states, “We had a great ride. R.I.P. 2000 - 2022,” and directs visitors to Screen9. This is current first-party evidence that the PicSearch service represented by the homepage is no longer operating under the historical service page; it is not evidence about every client that may still send the old header.
This profile therefore uses legacy-label. A present-day request carrying psbot-image could be a historical client, a retained integration, a test fixture, a malicious spoof, or an unrelated tool. Do not infer active PicSearch crawling, image-indexing permission, current operator ownership, AI input, model training, data retention, or access to image assets from the string alone.
Because no current policy was verified, do not publish a psbot-specific rule as if it were an operator instruction. If logs show a current match and your site policy is to exclude it, use the exact observed token in a deliberate site-owner rule:
User-agent: psbot-image
Disallow: /
For selective access to public image paths after confirmation:
User-agent: psbot-image
Allow: /public-images/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /originals/
Disallow: /api/
These are site-owner controls, not recovered PicSearch instructions. Robots.txt is advisory and cannot secure private, licensed, or original-resolution assets; use authentication, authorization, signed URLs, transformations, and origin controls.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. No current PicSearch source in this review published a psbot-image IP range, reverse-DNS procedure, Crawl-delay, rate guidance, removal address, or verification workflow. The header is useful for triage but is not authentication.
The discontinued homepage is a lifecycle signal, not a source-verification method. Check source IP and DNS evidence independently, and do not create a trust exception because the token names a former brand. If a network operator claims a revived or migrated service, request current first-party documentation and record the response and date before allowing access.
Compare observed behavior with the historical image-discovery hypothesis without turning it into attribution. Public thumbnails, image metadata, sitemaps, and ordinary assets may be requested by many clients. Original-resolution files, licensed media, private galleries, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not PicSearch identity or downstream use.
Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish. A discontinued bot page and a 404 do not establish robots compliance. If no exact group exists, a global rule may affect unrelated clients and should be adopted only as an explicit site-wide decision.
Page-level directives can express indexing preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate a historical crawler or secure private image routes. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes image indexing, AI input, reference use, and model training, document each purpose separately rather than inferring permission from a retired brand.
WAF and Nginx remediation examples
When the source is unverified or the service is discontinued, use report-only logging and a narrow observation. Avoid broad PicSearch, image, or browser rules that can affect unrelated clients:
{
"description": "Observe historical psbot-image candidates",
"expression": "lower(http.user_agent) contains \"psbot-image\"",
"action": "log"
}
After investigating a current match and deciding to protect sensitive media, scope enforcement to routes and preserve the evidence supporting the rule:
map $http_user_agent $block_historical_psbot_private {
default 0;
~*psbot-image 1;
}
server {
location ~ ^/(admin|account|private|licensed|originals|internal|api)/ {
if ($block_historical_psbot_private) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof and may identify a stale integration rather than PicSearch. Do not invent an IP allowlist, reverse-DNS suffix, rate, training policy, or permanent trust exception. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, image transformation, signed URLs, caching, and anomaly detection at the edge or origin.
Start in report-only mode, review false positives, and restrict only after evidence supports the action. Test public thumbnails, image metadata, sitemaps, original assets, licensed files, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and authorization instead of using robots.txt as an access-control mechanism.
Review checklist
Search logs for the exact psbot-image substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. Treat the historical brand and 404 page as lifecycle evidence, not proof of current ownership.
Re-check PicSearch’s current homepage, the historical bot URL, robots.txt, and any successor-service documentation. The current homepage says the PicSearch service ended in 2022, and the bot URL returned 404. Keep this profile at legacy-label unless a current authoritative source documents a revived or migrated crawler.
Decide whether your objective is to preserve public image discovery, reduce crawl load, limit extraction, protect licensed or original assets, or prevent private access. Publish exact robots rules only after the current token is known; enforce private media routes with authentication and origin controls.
Review image indexing, AI input, reference use, and model-training decisions separately. Neither a historical User-Agent nor a discontinued service page establishes downstream permission or use. Do not claim successful blocking or verification from configuration alone; validate later logs and obtain a current operator response if a live match appears.
Record why the profile is legacy-label and date any future evidence. If a successor service identifies itself differently, create a separate evidence trail rather than silently upgrading psbot.
References
- PicSearch homepage — current first-party page stating the service ended in 2022 and directing users to Screen9; reviewed with the headless browser on 2026-08-25.
- Historical PicSearch bot page — registry-linked page; returned
404 Not Foundduring headless-browser review on 2026-08-25. - Crawler User Agents community registry — registry context for the psbot-image label; it does not authenticate current infrastructure.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.