← Bot Directory/facebookbot
Bot directory / ai-training

How to block facebookbot: Meta's AI Training Crawler

A complete technical reference for facebookbot, the web crawler used by Meta to collect data for training AI models like Llama.

AI Summary: The facebookbot crawler is operated by Meta (Facebook) to scrape web content for training their AI models, including the Llama series. To opt out of Meta's AI training datasets, you should add User-agent: facebookbot to your robots.txt file and block its User-Agent string at your edge firewall.

Role and policy boundary

Meta operates several crawlers, but facebookbot (often alongside facebookexternalhit) has increasingly been utilized to gather vast amounts of public web data to train their generative AI models, such as the open-weight Llama models.

It is important to distinguish the policy boundary here. While facebookexternalhit is traditionally used to generate link previews when users share URLs on Facebook or Instagram, facebookbot is more directly associated with broad data ingestion. Blocking facebookbot is a targeted way to signal that your content should not be used for AI training by Meta, though careful configuration is required to ensure you do not accidentally break social media sharing functionality.

Layered verification

To effectively manage how Meta interacts with your site, a layered approach is essential.

The primary mechanism is the robots.txt file. Meta has stated that they respect the facebookbot directive for AI training opt-outs. By explicitly disallowing this User-Agent, you provide a clear policy signal.

However, because social media crawlers are deeply integrated into web traffic, relying solely on robots.txt might not be sufficient if you want strict enforcement. Network-level blocking provides a hard guarantee. By filtering requests based on the facebookbot User-Agent string, you can prevent the crawler from accessing your server entirely. When implementing this, ensure your rules specifically target facebookbot and not facebookexternalhit, unless you also intend to disable link previews on Meta platforms.

WAF and Nginx remediation examples

To enforce the block at the server level, you can implement User-Agent matching rules.

For an Nginx configuration, use the following snippet to return a 403 Forbidden status for requests identifying as facebookbot:

configuration / code
if ($http_user_agent ~* (facebookbot)) {
    return 403;
}

If you are protecting your site with Cloudflare WAF, you can deploy a custom rule using this JSON expression to block the traffic before it reaches your origin:

configuration / code
{
  "action": "block",
  "expression": "(http.user_agent contains \"facebookbot\")",
  "description": "Block Meta facebookbot for AI training opt-out"
}

Review checklist

To ensure your content is protected from Meta's AI training ingestion without breaking social sharing, follow these steps:

  1. Add User-agent: facebookbot with a Disallow: / directive to your robots.txt.
  2. Implement WAF or Nginx rules specifically targeting the facebookbot User-Agent string.
  3. Test sharing a link from your site on Facebook or Instagram to confirm that facebookexternalhit is still able to generate previews (if desired).
  4. Simulate a request using the facebookbot User-Agent to verify that the server correctly rejects the connection.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.