Google-CloudVertexBot: Robots.txt & Crawl Policy Reference
Guide to Google-CloudVertexBot, Google's crawler for enterprise Vertex AI Search. Learn how to configure robots.txt and access controls.
AI Summary: Google-CloudVertexBot is a crawler used when site owners build enterprise search applications using Google's Vertex AI (now Gemini Enterprise Agent Platform). It crawls content specifically to ingest it into a customer's private Vertex AI data store.
Role and policy boundary
Unlike Googlebot (which indexes for public Google Search) or Google-Extended (which controls AI training access), Google-CloudVertexBot is an enterprise crawler. It only crawls your site if you (or someone authorized on your domain) configure a Vertex AI Agent to ingest your content. It acts on behalf of the site owner's explicit configuration.
A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting. Start with a dedicated group:
User-agent: Google-CloudVertexBot
Allow: /
Disallow: /staging/
Disallow: /internal/
To stop access for the entire site, use:
User-agent: Google-CloudVertexBot
Disallow: /
Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect assistant and training crawlers too, and it makes later audits harder because the source of the decision is less specific.
Layered verification
Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.
Google-CloudVertexBot respects standard robots.txt rules. If you are building a Vertex AI Search application and want to ingest specific directories, you must ensure this bot is allowed. If you do not use Vertex AI, you can block it, though it typically will not crawl your site unless explicitly configured to do so in a Google Cloud console.
The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why a bot is allowed while another is blocked, rather than returning one blended website score.
Page-level directives can still override the intended outcome for indexing:
<meta name="robots" content="noai, noimageai">
X-Robots-Tag: noai, noimageai
If a response uses these tags, the report marks the result as blocked or conflicting even if the crawler-specific robots group is permissive. This is especially important for canonical pages served through an edge cache where headers may differ from the origin response.
WAF and Nginx remediation examples
{
"description": "Observe google-cloudvertexbot candidates",
"expression": "lower(http.user_agent) contains \"google-cloudvertexbot\"",
"action": "log"
}
If you need to strictly restrict access to your staging environment while allowing Vertex AI to index it, you can configure your WAF or Nginx to only allow this specific User-Agent from known Google IPs (after verifying the reverse DNS).
if ($http_user_agent ~* "Google-CloudVertexBot") {
# Allow Vertex AI crawler
break;
}
Use your platform's actual middleware response pattern rather than copying this simplified example without review. Never place a secret, verification token, or internal policy identifier in a public response header.
Review checklist
Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.
Verify that the dedicated group appears before relying on a wildcard, test a representative public and private path, and compare live response headers with robots.txt. Keep the policy close to the content owner's intent and record whether the site wants discovery, citation, or no access at all.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.