Brightbot: Robots.txt & Crawl Policy Reference
Technical reference for Brightbot 1.0, Bright Data's data collection crawler. Learn how to distinguish robots.txt from collectors.txt and control data collection responsibly.
AI Summary:
Brightbot 1.0is Bright Data's data collection crawler and the main collection pipeline for Bright Data products and services. Bright Data documents a unique User-Agent and source-IP subnet, domain-health monitoring, caching to reduce repeated downloads, and the vendor-specificcollectors.txtworkflow for approved collection guidelines. Userobots.txtfor a general crawl preference andcollectors.txtwhen working with Bright Data's documented Webmaster Console process.
Role and policy boundary
Brightbot is not a general search engine crawler. Bright Data describes it as the data collection crawler used across its products and services. That means a request may support public-web datasets, research, market intelligence, or AI-oriented data workflows, but the public User-Agent alone does not tell a site owner the exact downstream product or customer use case.
Bright Data provides additional governance controls around this crawler. Its documentation describes a Webmaster Console, a unique User-Agent and source-IP subnet, a cache layer intended to prevent repetitive downloads within a 24-hour period, and health monitoring that can apply rate limits when traffic affects a domain. It also documents collectors.txt, a Bright Data-specific resource for communicating endpoint restrictions, PII boundaries, interactive-element exclusions, and copyright-related guidance. collectors.txt is useful in that vendor relationship, but it is not a replacement for authentication and it is not a universal web standard.
A site may allow Brightbot on public, low-risk pages while excluding login, account-management, transactional, ad, review, and customer-data endpoints. If the objective is to opt out of third-party collection entirely, publish a dedicated robots rule and enforce the decision at the edge as needed:
User-agent: Brightbot
Disallow: /
For a selective policy, use an explicit allowlist of public paths:
User-agent: Brightbot
Allow: /docs/
Allow: /public-data/
Disallow: /staging/
Disallow: /internal/
Disallow: /account/
A robots rule is a published crawl preference; it does not authenticate a caller, protect a private endpoint, or guarantee how a third-party crawler will use data. Keep private information behind real authorization controls.
Layered verification
Verify Brightbot through independent signals. Start with the complete User-Agent, including the documented Brightbot 1.0 value, and compare the source address with Bright Data's current source-IP information or the Web Master Console where available. A User-Agent can be copied, so it is not sufficient on its own for an allowlist or sensitive-data decision.
Next, check the top-level /robots.txt response and the exact path scope. If your organization is using Bright Data's workflow, review the collectors.txt guidelines in the Webmaster Console and confirm that the requested endpoint is within the approved collection boundary. These controls answer different questions: robots.txt communicates a general crawler preference, while collectors.txt communicates Bright Data-specific collection rules and restrictions.
Bright Data also describes health monitoring and rate limiting as protective technology. That is a vendor-side mitigation and does not mean that a site owner can assume every request is harmless. Monitor request volume, latency, status codes, cache behavior, and access to interactive endpoints in your own logs.
The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. An allowed robots result therefore means that the published preference permits the request; it does not establish the caller's identity, validate a source subnet, or grant access to a protected route.
Page-level metadata is a separate signal and is not a substitute for blocking private content:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
Use these directives when the concern is indexing or discoverability, and use authentication, authorization, WAF rules, and rate limits when the concern is access or resource consumption.
WAF and Nginx remediation examples
If you decide to block the declared Brightbot token, match the stable identifier rather than the entire string. This example blocks self-declared Brightbot traffic and should be combined with source validation and rate limiting if you need a stronger identity decision:
{
"description": "Block declared Brightbot crawler",
"expression": "lower(http.user_agent) contains \"brightbot\"",
"action": "block"
}
For a selective Nginx policy, protect only the paths that should not be collected:
map $http_user_agent $deny_brightbot {
default 0;
~*brightbot 1;
}
server {
location /account/ {
if ($deny_brightbot) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not use a User-Agent match as the only control for confidential data, and do not treat collectors.txt as an access-control protocol. If you participate in Bright Data's Webmaster Console process, document the approved collection boundary and compare it with application logs. If you do not, use the standard robots, authentication, WAF, and rate-limit layers appropriate to your risk.
Review checklist
Request /robots.txt from the canonical host and verify the exact Brightbot group, status, content type, and path scope. Test one public documentation path, one disallowed internal path, and one interactive or account path. Record User-Agent, source IP, HTTP status, final URL, redirect chain, response size, and request rate.
If Bright Data is an approved collector, review collectors.txt through the documented Webmaster Console workflow and confirm that PII, interactive endpoints, and private routes are excluded as intended. Re-test after a CDN, WAF, origin, or policy change. Keep the distinction explicit in the audit: vendor controls can reduce abuse, but they do not remove the site's responsibility to protect private data.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.