Bot directory / ai-search

TavilyBot: Robots.txt & Crawl Policy Reference

Technical reference for the TavilyBot registry label associated with Tavily's web-access APIs. Learn how to verify actual traffic without inferring a crawler contract from API documentation.

AI Summary: Tavily's official site and docs describe a web-access layer for AI agents with Search, Extract, Crawl, Map, and Research APIs that retrieve live web data. They do not publish a dedicated TavilyBot crawler contract in the pages reviewed, so the registry token, source ranges, crawl rate, and robots behavior remain unverified. Treat the header as an observation and enforce unwanted access with layered controls.

Role and policy boundary

Tavily officially presents its platform as a web-access layer for agents. Its documentation covers searching the web, extracting webpages, crawling sites, mapping links, and creating research tasks; returned content is structured and chunked so models can reason over fresh sources. This establishes API and product capabilities, but it does not prove that Tavily operates a centralized crawler called TavilyBot against every site.

The registry lists Mozilla/5.0 (compatible; TavilyBot; +https://tavily.com/). The official product and documentation pages reviewed for this profile do not publish a dedicated User-Agent specification, source IP list, crawl schedule, robots behavior, or webmaster verification process. An API customer may also retrieve a URL through Tavily without the site receiving a request that declares this exact token.

Treat TavilyBot as a partially documented product-associated label. Blocking it may affect one possible Tavily access path, but it does not remove pages from ordinary search engines or control other AI systems. Allowing it does not grant permission to access private, paywalled, customer-specific, or authenticated content.

If your logs confirm the exact token and you want to communicate a restriction, publish:

configuration / code
User-agent: TavilyBot
Disallow: /

For a selective public-reference policy:

configuration / code
User-agent: TavilyBot
Allow: /docs/
Allow: /public-reference/
Disallow: /private/
Disallow: /account/
Disallow: /api/

Robots.txt is advisory. Protect confidential content with authentication and application authorization.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. The browser-like token is self-declared, and the reviewed Tavily docs publish no IP list with which to authenticate it.

Analyze behavior without assigning a purpose prematurely. Search or extract requests may target a single user-selected URL, while crawl or map operations may traverse several pages. High concurrency, repeated retries, media downloads, or requests for private APIs are operational signals, not proof of a particular Tavily product or model-training use.

Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact group, and path match. Because the official docs reviewed do not establish TavilyBot's robots behavior, do not mark a request compliant solely because a robots rule allowed it. Page-level directives may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not protect private routes and may not be honored by an unverified client. Use authentication, signed URLs, and origin controls for sensitive content.

WAF and Nginx remediation examples

If logs confirm unwanted requests carrying the token, use a narrow WAF rule:

configuration / code
{
  "description": "Block observed TavilyBot token",
  "expression": "lower(http.user_agent) contains \"tavilybot\"",
  "action": "block"
}

For Nginx, scope enforcement to high-risk routes while investigating public access:

configuration / code
map $http_user_agent $block_tavilybot {
    default 0;
    ~*TavilyBot 1;
}

server {
    location ~ ^/(private|internal|account|paywall|api)/ {
        if ($block_tavilybot) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof or evade and may block an approved Tavily integration. Do not create an IP allowlist without a verified operator-published range. Test browsers, feed readers, social previews, search crawlers, and internal monitors. Combine the rule with route authorization, rate limits, signed URLs, and anomaly detection.

Review checklist

Search logs for TavilyBot and preserve source network, paths, response sizes, status, timing, and rate. Re-check Tavily's current product and documentation pages before treating the registry label as a confirmed crawler. During this review the official pages documented web-access APIs but not a dedicated TavilyBot identity or policy.

Decide whether your objective is to preserve potential AI-search access, prevent possible extraction, protect private content, or reduce crawl load. Publish a targeted robots group for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, feeds, sitemaps, media, uploads, and APIs separately. Revisit the profile if Tavily publishes a canonical crawler User-Agent, source verification, rate policy, or robots statement.

References

  1. Tavily — official product page describing real-time web access for agents.
  2. Tavily Documentation — official Search, Extract, Crawl, Map, and Research API documentation; no dedicated TavilyBot policy was found on the reviewed page.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.