Meta Robots Tag
An HTML element placed in the document head that provides page-level instructions to search engine crawlers regarding indexing, link following, and snippet rendering.
AI Summary: The meta robots tag is an HTML directive located in the document head that controls how search engines and AI crawlers index and display a specific page. Common directives include noindex, nofollow, and noarchive, which override broader robots.txt rules.
Technical Definition
The Meta Robots tag is an HTML element embedded in the <head> section of a web page:
<meta name="robots" content="noindex, nofollow, noarchive">
It gives publishers granular, page-level control over search engines and automated crawlers, superseding broader path-based rules declared in /robots.txt.
Critical Directives for AI Governance
noindex: Instructs crawlers not to include the page in public search indexes.noarchive: Forbids engines from saving cached versions of the page.max-snippet:[number]: Limits the number of characters displayed in text snippets.nosnippet: Prohibits search engines and generative models from displaying text extracts.
Important Rule: The Robots.txt Conflict
[!WARNING] If a page is blocked via
robots.txt(Disallow: /page), the crawler will never fetch the page HTML. Consequently, it will never see the<meta name="robots" content="noindex">tag! To remove an existing page from search results, allow it inrobots.txtwhile serving thenoindextag.
Detect conflicting robots directives across your web properties. Audit with Geolify.ai.