Gemini-Deep-Research: Robots.txt & Crawl Policy Reference
Technical reference for Google's Gemini-Deep-Research crawler. Learn how to manage traffic from this agentic research tool using robots.txt and WAF rules.
AI Summary:
Gemini-Deep-Researchis a User-Agent token used by Google's Gemini platform for its agentic Deep Research feature. This tool autonomously browses websites to synthesize information for user-directed research tasks. It is distinct fromGooglebot(search indexing) andGoogle-Extended(model training). Site owners can manage this traffic using standardrobots.txtdirectives or network-level rules.
Role and policy boundary
Google's Gemini Deep Research is an advanced, multi-step research agent. When a user requests deep research on a topic, the agent autonomously navigates the web, reads content, and synthesizes findings. The Gemini-Deep-Research User-Agent identifies requests made during these active research sessions.
It is important to understand the boundary between this tool and other Google crawlers:
- Googlebot: Indexes content for Google Search.
- Google-Extended: Controls whether content is used to improve Gemini foundation models.
- Gemini-Deep-Research: Fetches live content specifically to answer a user's immediate, complex query.
Blocking Gemini-Deep-Research prevents Gemini users from using your site as a source during their active research tasks, but it does not prevent your site from appearing in standard Google Search results, nor does it automatically opt you out of foundation model training.
If you wish to prevent this agent from accessing your site, you can declare a specific robots rule:
User-agent: Gemini-Deep-Research
Disallow: /
Layered verification
Because this crawler originates from Google, you can verify its authenticity using reverse DNS lookups, similar to how you verify Googlebot.
- Log Analysis: Look for the
Gemini-Deep-Researchtoken in the User-Agent string. - IP Verification: Perform a reverse DNS lookup on the accessing IP address. It should resolve to a Google domain (e.g.,
googlebot.comorgoogle.com). Then, perform a forward DNS lookup on that domain name to ensure it matches the original IP. - Behavioral Patterns: Expect this traffic to occur in bursts corresponding to user research sessions, rather than the steady, continuous crawl pattern of a search indexer.
While robots.txt is the primary control mechanism, you can also use HTML metadata to manage indexing preferences, though an active research agent may still read the page to answer a user query even if it is marked noindex.
<meta name="robots" content="noindex, nofollow">
WAF and Nginx remediation examples
If you want to enforce a block at the network edge, ensuring that even spoofed requests are stopped before reaching your application, you can use a WAF or Nginx rule.
For a WAF (like Cloudflare), you can match the specific token:
{
"description": "Block Gemini Deep Research",
"expression": "lower(http.user_agent) contains \"gemini-deep-research\"",
"action": "block"
}
To implement this block in Nginx:
map $http_user_agent $block_gemini_research {
default 0;
~*Gemini-Deep-Research 1;
}
server {
if ($block_gemini_research) {
return 403;
}
# Application routing continues here
}
Note: If you only want to block unauthorized spoofing, combine the User-Agent check with an IP range or ASN check to allow verified Google traffic while blocking imposters.
Review checklist
To manage traffic from the Gemini Deep Research agent:
- Define Policy: Decide whether you want your content to be accessible to users performing deep research via Gemini.
- Implement Robots.txt: Add a
User-agent: Gemini-Deep-Researchrule to your/robots.txtfile to declare your preference. - Verify Traffic: Use reverse DNS to confirm that requests claiming this User-Agent actually originate from Google infrastructure.
- Deploy Edge Rules: If you require strict enforcement, implement WAF or Nginx rules to block the User-Agent, optionally allowing verified Google IPs if you want to permit the official crawler while blocking spoofers.
- Monitor Logs: Periodically review access logs to ensure the policy is working as intended and that the crawler respects your directives.
References
- Gemini Deep Research Overview — Google's introduction to the agentic research feature.
- Google Crawlers Documentation — Official documentation for managing various Google crawlers and agents.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.