← Bot Directory/FirecrawlAgent
Bot directory / ai-training

FirecrawlAgent: Robots.txt & Crawl Policy Reference

Technical reference for Firecrawl, a web scraping and context extraction API for AI agents. Learn how to manage its default crawler and handle user-configured instances.

AI Summary: Firecrawl is a web scraping, crawling, and extraction API designed to feed live web context into AI agents. By default, it identifies itself as FirecrawlAgent and checks robots.txt. However, because it is a configurable tool (both hosted and open-source), operators can override the User-Agent and bypass robots rules. Site owners must combine robots directives with WAF rules and behavioral monitoring to effectively manage this traffic.

Role and policy boundary

Firecrawl is not a traditional search engine indexer or a single centralized foundation model crawler. It is an infrastructure tool used by developers to build AI agents that need to search, scrape, and interact with websites. It can convert entire websites into clean Markdown or structured data.

According to Firecrawl's developers, the default configuration respects robots.txt directives targeting the FirecrawlAgent token (or the * wildcard). However, the API provides explicit options for users to override headers (including the User-Agent) and to ignore robots.txt entirely. This means that while a robots rule will stop the default, compliant instances, it will not stop an operator who has chosen to spoof a standard web browser or bypass crawl restrictions.

For site owners, the policy boundary depends on whether you want to allow AI agents to extract your content. If you want to block default Firecrawl behavior, a robots rule is the first step. If you want to protect your site from aggressive or disguised scraping using the Firecrawl engine, you must implement active network defenses.

To block the default, compliant Firecrawl crawler, add this to your /robots.txt:

configuration / code
User-agent: FirecrawlAgent
Disallow: /

Layered verification

Because Firecrawl can be run by anyone and configured to use any User-Agent, relying solely on the FirecrawlAgent token is insufficient for comprehensive protection.

Start by verifying logs for the default token. If you see FirecrawlAgent, the operator is likely using the default settings and will honor your robots rules.

If you suspect disguised Firecrawl traffic, you must look beyond the User-Agent:

  1. Behavioral Patterns: Look for rapid, parallel requests fetching HTML content without loading associated assets (CSS, JS, images) in a way a real browser would, unless the operator is using Firecrawl's advanced browser rendering features.
  2. IP Sources: Hosted Firecrawl API requests may originate from specific cloud provider IP ranges. Open-source instances will originate from wherever the operator has deployed them.
  3. Session Consistency: Disguised scrapers often fail to maintain consistent session state, cookies, or TLS fingerprints compared to the browsers they claim to be.

HTML metadata like <meta name="robots" content="noindex"> is intended for search engines. A direct scraping tool like Firecrawl, commanded by a user to extract a specific page, will generally ignore indexing directives and extract the content anyway.

WAF and Nginx remediation examples

To block the declared FirecrawlAgent token at the network edge, use a WAF rule. This prevents the request from reaching your application, saving resources.

For a standard WAF configuration:

configuration / code
{
  "description": "Block default FirecrawlAgent",
  "expression": "lower(http.user_agent) contains \"firecrawlagent\"",
  "action": "block"
}

To enforce this block in Nginx:

configuration / code
map $http_user_agent $block_firecrawl {
    default 0;
    ~*FirecrawlAgent 1;
}

server {
    if ($block_firecrawl) {
        return 403;
    }
    
    # Application routing continues here
}

To stop instances of Firecrawl that have been configured to use a spoofed User-Agent, you must deploy advanced bot management features provided by your WAF (such as Cloudflare Bot Management or AWS WAF Bot Control). These systems evaluate TLS fingerprints, JavaScript execution challenges, and request anomalies to block automated tools regardless of their declared identity.

Review checklist

To manage Firecrawl and similar agentic extraction tools:

  1. Implement Robots Policy: Add FirecrawlAgent to your /robots.txt to block operators using default settings.
  2. Monitor Logs: Search your access logs for FirecrawlAgent to gauge the volume of compliant scraping attempts.
  3. Deploy Edge Blocking: Implement WAF or Nginx rules to block the declared token, reducing load on your origin servers.
  4. Enable Advanced Bot Protection: Turn on behavioral and challenge-based bot protection in your WAF to catch Firecrawl instances that override their User-Agent.
  5. Protect Sensitive Routes: Ensure that all private, transactional, or high-cost API routes require strong authentication, as scraping tools can easily bypass unauthenticated endpoints.

References

  1. Firecrawl Documentation — Official documentation detailing scraping, crawling, and API configuration.
  2. Firecrawl GitHub Repository — Open-source repository where issues confirm the configurable nature of the User-Agent and robots compliance.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.