Robots.txt Protocol
A plain text configuration file placed at the root of a domain that establishes declarative access policies for automated web crawlers according to RFC 9309.
AI Summary: The robots.txt file is a publicly accessible text document governed by RFC 9309 that declares which parts of a site automated crawlers are allowed or forbidden to access. It serves as the primary declarative interface between publishers and AI crawlers.
Technical Definition
The robots.txt file is a plain text document stored at the root directory of a domain (e.g., https://example.com/robots.txt). Formally standardized in RFC 9309 (Robots Exclusion Protocol), it provides declarative guidelines to automated user-agents about which URL paths they are authorized to request.
Core Directives
User-agent: Specifies which bot the subsequent rules apply to.Disallow: Designates a path prefix that the crawler must not fetch.Allow: Overrides a broad Disallow rule for a specific sub-path.Sitemap: Declares the location of the site's XML sitemap index.
Best-Practice AI Governance Template
# 1. Broad allow for standard search engines
User-agent: *
Disallow: /admin/
Disallow: /private/
Sitemap: https://example.com/sitemap.xml
# 2. Block general AI training harvesting
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
# 3. Allow AI search for citation generation
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
Check your live robots.txt for syntax conflicts, wildcard bugs, and AI policy mismatches. Audit with Geolify.ai.