← Bot Directory/img2dataset
Bot directory / ai-training

img2dataset: Robots.txt & Crawl Policy Reference

Technical reference for img2dataset, an open-source image downloader used to build machine-learning datasets. Learn why its identity and compliance depend on the operator.

AI Summary: img2dataset is an open-source command-line tool that downloads image URLs and converts them into machine-learning datasets. It is not one centralized crawler: anyone can run and configure it. The registry User-Agent is an observed example, not proof of origin or a universal default. Site owners should communicate policy with robots.txt, then enforce protection with rate limits, WAF controls, signed URLs, and access authorization.

Role and policy boundary

The official img2dataset repository describes a tool for turning large collections of image URLs into datasets. It can download, resize, and package images at large scale. That makes it a useful component in data engineering and computer-vision workflows, but it does not represent a single public search or model-training service.

The distinction between software and operator is essential. An organization may run img2dataset to collect a permitted dataset, while another operator may use it to download public images without permission. The open-source repository cannot establish the purpose, legal basis, or policy of each deployment. A registry-listed User-Agent can help identify an operator that chooses to send it, but the header is self-declared and configurable.

If your logs confirm that the observed token is being used and you want to communicate that images should not be fetched, publish:

configuration / code
User-agent: img2dataset
Disallow: /

For a selective policy that exposes low-resolution public assets but not originals:

configuration / code
User-agent: img2dataset
Allow: /public-thumbnails/
Disallow: /originals/
Disallow: /uploads/
Disallow: /private-media/
Disallow: /api/

Because img2dataset can be configured or forked, a robots rule is not a guarantee against every instance. Private or commercially restricted images need authorization, storage controls, and edge enforcement.

Layered verification

Start by examining access logs for the complete observed header, including the img2dataset token and repository URL when present. Record source IP, ASN, reverse DNS, timestamp, request path, method, status, response size, cache result, and request rate. Do not treat a matching User-Agent as identity authentication.

Analyze the traffic pattern. Dataset downloads often request many image URLs, repeat retries, omit ordinary browser asset behavior, and concentrate on high-resolution or original files. These signs can establish bandwidth and resource risk, but they do not prove that the operator is using the upstream repository or that the collection is for model training.

No centralized img2dataset source range or universal robots contract exists because the software is operator-run. Evaluate /robots.txt as a published preference and test the exact path matching. Page directives can communicate discoverability preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

For image protection, do not rely on these directives. Use object permissions, signed URLs with expiry, origin shielding, and request authentication when the asset must not be downloaded anonymously.

WAF and Nginx remediation examples

If you have confirmed unwanted requests declaring img2dataset, match the token at the WAF:

configuration / code
{
  "description": "Block declared img2dataset downloader",
  "expression": "lower(http.user_agent) contains \"img2dataset\"",
  "action": "block"
}

For Nginx, scope the rule to media and API routes first:

configuration / code
map $http_user_agent $block_img2dataset {
    default 0;
    ~*img2dataset 1;
}

server {
    location ~ ^/(images|media|uploads|originals|api)/ {
        if ($block_img2dataset) { return 403; }
        try_files $uri $uri/ =404;
    }
}

This rule catches only the declared token. An operator can replace it with a browser-like header, and an overly broad rule can block a legitimate internal dataset job. Combine edge matching with rate limits, signed asset URLs, hotlink restrictions, and authentication. Test against browsers, social preview fetchers, image optimizers, and internal jobs before enforcing globally.

Review checklist

Confirm whether the traffic actually declares img2dataset; do not infer it from high-volume images alone. Preserve the full request and determine the affected paths, response sizes, concurrency, and source network. Decide whether your objective is to prevent dataset collection, reduce bandwidth, protect private images, or control indexing.

Publish a specific robots group when it helps communicate your preference, but treat it as advisory for operator-run software. Deploy WAF or Nginx enforcement for unwanted declared traffic, and protect sensitive media with signed URLs or authorization. Test thumbnails, originals, uploads, image sitemaps, and API endpoints independently. Revisit the policy when the deployment, User-Agent, or business purpose changes.

References

  1. img2dataset repository — official open-source project description and implementation source.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.