Google-NotebookLM: Robots.txt & Crawl Policy Reference
Technical reference for Google-NotebookLM, a user-triggered fetcher. Learn why it ignores robots.txt and how to manage this traffic using network rules.
AI Summary:
Google-NotebookLMis an official User-Agent token for Google's NotebookLM (also known as Gemini Notebook). Unlike standard search crawlers, it is classified as a "user-triggered fetcher." This means it fetches a webpage only when a user explicitly requests to import that specific URL into their private notebook. Because it acts on behalf of a direct user request, it ignoresrobots.txtrules. To block this tool, you must use WAF or Nginx rules.
Role and policy boundary
Google NotebookLM is an AI-powered research and note-taking application. Users can upload documents or provide URLs, and the tool will read the content to generate summaries, study guides, and audio overviews (like podcasts) based only on the provided sources.
When a user provides a URL to NotebookLM, Google's servers fetch the page using the Google-NotebookLM User-Agent. Google classifies this as a user-triggered fetcher. This classification is crucial for policy boundaries:
- Not a Search Indexer: It does not crawl your site to add pages to Google Search.
- Not a Training Crawler: Fetching a page for a user's notebook does not automatically add that content to the training dataset for Gemini foundation models (which is governed by
Google-Extended). - Ignores Robots.txt: Because the fetch is initiated by a user acting as if they were opening the page in a browser, Google's policy for user-triggered fetchers is to bypass
robots.txt. ADisallowrule will not stop NotebookLM from reading the page.
If you want to prevent users from importing your content into NotebookLM, you cannot rely on standard robots directives. You must actively block the User-Agent at the network level.
Layered verification
To verify that traffic claiming to be Google-NotebookLM is legitimate:
- Log Analysis: Check your server logs for the exact
Google-NotebookLMtoken. - Reverse DNS: Because it is a Google service, the IP address should resolve to a Google domain (e.g.,
google.com). Perform a reverse DNS lookup, followed by a forward DNS lookup on the result, to confirm the IP belongs to Google. - Request Pattern: Traffic will be sporadic and directly correlated with individual users attempting to import specific pages, rather than a systematic crawl of your entire site hierarchy.
HTML indexing directives like <meta name="robots" content="noindex"> are ineffective here. The fetcher is not indexing the page for public search; it is retrieving the text for a private user session.
WAF and Nginx remediation examples
Because robots.txt is ignored, you must use application or network-level blocking to stop NotebookLM from fetching your content.
Using a WAF (such as Cloudflare), you can create a rule that blocks the specific User-Agent:
{
"description": "Block Google-NotebookLM fetcher",
"expression": "lower(http.user_agent) contains \"google-notebooklm\"",
"action": "block"
}
If you manage your own Nginx server, you can block the request before it reaches your application logic:
map $http_user_agent $block_notebooklm {
default 0;
~*Google-NotebookLM 1;
}
server {
if ($block_notebooklm) {
# Return 403 Forbidden to block the fetch
return 403;
}
# Standard application configuration
}
Note: Blocking this User-Agent will result in an error message for the NotebookLM user attempting to import the URL, stating that the content could not be accessed.
Review checklist
To manage access from Google NotebookLM:
- Understand the Tool: Recognize that NotebookLM is a private research tool, not a public search indexer or a broad training crawler.
- Acknowledge Robots.txt Limitations: Accept that adding
Google-NotebookLMto yourrobots.txtwill not prevent the tool from fetching your pages. - Implement Network Blocks: If you decide to restrict access, deploy WAF or Nginx rules to block the
Google-NotebookLMUser-Agent string. - Verify IP Addresses: Use reverse DNS to ensure you are blocking legitimate Google traffic and not spoofed requests, though a broad User-Agent block will stop both.
- Monitor Impact: Review logs to see how often users attempt to import your content into NotebookLM and adjust your policies if necessary.
References
- Google User-Triggered Fetchers — Official documentation explaining the role of user-triggered fetchers.
- NotebookLM (Gemini Notebook) — The official product page for the research tool.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.