← Glossary/Robots.txt Protocol
Glossary Term

Robots.txt Protocol

A plain text configuration file placed at the root of a domain that establishes declarative access policies for automated web crawlers according to RFC 9309.

AI Summary: The robots.txt file is a publicly accessible text document governed by RFC 9309 that declares which parts of a site automated crawlers are allowed or forbidden to access. It serves as the primary declarative interface between publishers and AI crawlers.

Technical Definition

The robots.txt file is a plain text document stored at the root directory of a domain (e.g., https://example.com/robots.txt). Formally standardized in RFC 9309 (Robots Exclusion Protocol), it provides declarative guidelines to automated user-agents about which URL paths they are authorized to request.

Core Directives

  • User-agent: Specifies which bot the subsequent rules apply to.
  • Disallow: Designates a path prefix that the crawler must not fetch.
  • Allow: Overrides a broad Disallow rule for a specific sub-path.
  • Sitemap: Declares the location of the site's XML sitemap index.

Best-Practice AI Governance Template

configuration / code
# 1. Broad allow for standard search engines
User-agent: *
Disallow: /admin/
Disallow: /private/
Sitemap: https://example.com/sitemap.xml

# 2. Block general AI training harvesting
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

# 3. Allow AI search for citation generation
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Check your live robots.txt for syntax conflicts, wildcard bugs, and AI policy mismatches. Audit with Geolify.ai.

Related terms