← Bot Directory/ImagesiftBot
Bot directory / search-engine

ImagesiftBot: Robots.txt & Crawl Policy Reference

Technical reference for ImagesiftBot, including its published image-indexing purpose, robots.txt and Crawl-delay behavior, stored fields, and opt-out contact.

AI Summary: ImageSift documents ImagesiftBot as a crawler for publicly available images used in web-intelligence products and similar-image search. It publishes the exact User-Agent Mozilla/5.0 (compatible; ImagesiftBot; +imagesift.com), says targeted robots.txt rules and Crawl-delay are respected, and provides support@imagesift.com for opt-out requests. No IP range or reverse-DNS verification method was published on the reviewed page.

Role and policy boundary

ImageSift’s official page identifies ImagesiftBot as a web crawler that scrapes the internet for publicly available images to support its web-intelligence products. It states that the crawler saves the host URL, page text, and image alt text alongside images, then stores the information in an index used by ImageSift products for search and retrieval of similar images.

The published User-Agent is:

configuration / code
Mozilla/5.0 (compatible; ImagesiftBot; +imagesift.com)

This is a documented image-discovery and indexing purpose. It does not establish permission to access private or licensed media, authenticated image URLs, account areas, or origin-only endpoints. Robots.txt is the site-owner mechanism described by ImageSift for crawl preferences; it is not a substitute for authentication or authorization.

ImageSift says that standard directives targeting ImagesiftBot are respected. Its example permits all pages except /private/:

configuration / code
User-Agent: ImagesiftBot
Allow: /
Disallow: /private/

It also documents Crawl-delay as the minimum duration, in seconds, between the starts of consecutive requests. For example:

configuration / code
User-Agent: ImagesiftBot
Crawl-delay: 5

The page says this divides each day into five-second intervals and allows at most one request to the domain inside each interval. If there is no ImagesiftBot-specific rule but a Googlebot rule exists, ImageSift says ImagesiftBot follows the Googlebot directives. Treat that inheritance as a documented behavior of the ImageSift service, not as evidence that ImagesiftBot is Googlebot or that Google infrastructure is involved.

To restrict the bot from private paths:

configuration / code
User-agent: ImagesiftBot
Allow: /public/
Allow: /images/public/
Disallow: /private/
Disallow: /uploads/originals/
Disallow: /account/
Disallow: /api/
Crawl-delay: 10

For a complete exclusion:

configuration / code
User-agent: ImagesiftBot
Disallow: /

ImageSift lists support@imagesift.com for questions and opt-out requests. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. The official page publishes the User-Agent and policy behavior but does not publish an IP range or reverse-DNS verification procedure. Therefore, the header is useful for triage but is not authentication.

Compare observed behavior with the documented image-indexing role. Requests for public image assets, image URLs, page text, and alt-text-bearing HTML may be consistent with ImageSift’s stated purpose. Private endpoint access, high concurrency, repeated retries, unexpected non-image downloads, or traffic from an unverified source may indicate spoofing, abuse, a changed deployment, or another client. These observations establish impact, not operator identity or downstream use.

Evaluate /robots.txt independently. Confirm that the response is served by the correct host, returns a successful status and text content type, contains an exact ImagesiftBot group, and applies Allow, Disallow, and Crawl-delay as intended. If you rely on fallback Googlebot rules, test the parser and document that choice explicitly. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, noimageindex">
configuration / code
X-Robots-Tag: noindex, noimageindex

These signals do not secure private routes or authenticate the crawler. Enforce sensitive boundaries in the application and at the origin, and do not assume that an image-level directive removes previously indexed data without verifying the service’s response.

WAF and Nginx remediation examples

If you want to observe the crawler before enforcing a policy, begin with a narrow report-only rule. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe ImagesiftBot image-indexing requests",
  "expression": "lower(http.user_agent) contains \"imagesiftbot\"",
  "action": "log"
}

After reviewing false positives and confirming your robots decision, scope enforcement to private or expensive routes:

configuration / code
map $http_user_agent $block_imagesift_private {
    default 0;
    ~*ImagesiftBot 1;
}

server {
    location ~ ^/(private|account|uploads/originals|internal|api)/ {
        if ($block_imagesift_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and can catch an authorized image-monitoring integration if it is too broad. Do not invent an IP allowlist or reverse-DNS suffix because none was published on the reviewed page. Test image assets, HTML containing alt text, thumbnails, original uploads, manifests, feeds, sitemaps, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, hotlink controls, and anomaly detection.

Review checklist

Search logs for the complete published header and record source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Confirm that the traffic is limited to public image and page resources when that is the expected purpose; do not classify it as authentic from the header alone.

Check that your exact ImagesiftBot robots group expresses the intended search-visibility and image-licensing decision. Test Crawl-delay with the effective request intervals, and verify whether a Googlebot fallback rule is intentionally being applied. Protect private and original-media routes with application authorization, not robots.txt.

If you want to opt out, send the domain and representative evidence to support@imagesift.com, then verify subsequent access logs and any operational response. Re-check the official page for changes to the User-Agent, stored fields, fallback behavior, or opt-out contact. Do not claim successful blocking from configuration alone.

The documented ImageSift page describes image and associated page metadata indexing; it does not establish AI-model training use. Keep that boundary explicit when interpreting logs or communicating with site owners.

References

  1. ImageSift Bot documentation — official page documenting purpose, exact User-Agent, stored host/text/alt-text fields, robots behavior, Crawl-delay, Googlebot fallback, and opt-out contact; reviewed with the headless browser on 2026-08-25.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.