Googlebot-Image: Robots.txt & Crawl Policy Reference
Technical reference for Googlebot-Image, Google's dedicated crawler for indexing images in Google Search.
AI Summary:
Googlebot-Imageis the official web crawler used by Google specifically to discover and index images for Google Image Search. It fully respectsrobots.txtdirectives targeted at theGooglebot-Imagetoken. Google provides robust verification methods, including a published IP list and reverse-DNS lookups, allowing site owners to authenticate traffic and safely manage image indexing.
Role and policy boundary
Google operates Googlebot-Image to crawl the web specifically for images. When Google discovers an image on a webpage, it uses this crawler to fetch the image file itself.
According to Google's official documentation, Googlebot-Image respects robots.txt. If you want to allow Google to index your HTML pages but prevent your images from appearing in Google Image Search, you can block Googlebot-Image while allowing Googlebot.
To exclude your images from Google Image Search across your entire site, add the following to your robots.txt:
User-agent: Googlebot-Image
Disallow: /
For selective access (e.g., preventing the indexing of specific image directories):
User-agent: Googlebot-Image
Disallow: /private-images/
Disallow: /licensed-assets/
These are explanatory site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls. Blocking Googlebot-Image in robots.txt prevents Google from crawling the image files, which effectively removes them from Google Images.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate.
Google officially supports two methods to verify Googlebot-Image traffic:
- IP Allowlist: Google publishes a JSON file containing all active Googlebot IP ranges.
- Reverse DNS: Perform a reverse DNS lookup on the accessing IP address. Verify that the hostname ends with
.googlebot.comor.google.com. Then, perform a forward DNS lookup on that hostname to confirm it resolves back to the original IP address.
Do not treat the word Googlebot-Image in a header as proof of ownership. Malicious actors frequently spoof Googlebot User-Agents to bypass security controls. Check source IP and DNS evidence independently before creating a trust exception.
Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish.
Page-level directives can express indexing preferences for images:
<meta name="robots" content="noimageindex">
X-Robots-Tag: noimageindex
These signals do not authenticate a crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
When configuring WAF rules, always pair User-Agent matching with IP verification. Avoid broad googlebot rules that blindly trust the header:
{
"description": "Observe unverified Googlebot-Image candidates",
"expression": "lower(http.user_agent) contains \"googlebot-image\"",
"action": "log"
}
After independent verification (using Google's IP list or reverse-DNS) and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule:
map $http_user_agent $is_googlebot_image {
default 0;
~*Googlebot-Image 1;
}
# Note: In a real Nginx configuration, you must pair this with IP validation
# (e.g., using the geo module with Google's published IP list) to prevent spoofing.
server {
location ~ ^/(admin|account|private|licensed|internal)/ {
# Only allow verified Google IPs here, or block entirely
if ($is_googlebot_image) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.
Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.
Review checklist
Search logs for the exact Googlebot-Image substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.
Re-check the crawler registry, Google's official documentation, robots.txt, and operator contact path. Verify the IPs using Google's published list or the .googlebot.com reverse-DNS method.
Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.
Review search indexing, AI input, reference use, and model-training decisions separately. Googlebot-Image is explicitly for image search indexing. Do not claim successful blocking or verification from configuration alone; validate later logs.
Record the date and reason for the documented classification so future evidence can be compared. If a new deployment identifies itself differently, create a separate evidence trail.
References
- Google Search Central: Googlebot — official documentation on Googlebot and Googlebot-Image.
- Verifying Googlebot — instructions for verifying Google crawlers via IP or DNS.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.