atlassian-bot: Robots.txt & Crawl Policy Reference
Technical reference for the Atlassian Rovo web crawler associated with custom website search and indexing. Learn how to expose or restrict public content safely.
AI Summary:
atlassian-botis the User-Agent label associated with Atlassian's web crawling for custom website connections to Teamwork Graph and Rovo. Atlassian documents that this connector performs a search crawl and index so eligible content can appear in Atlassian Search and be used by Rovo Chat and Agents. The exact public bot contract is only partially documented, so treat the User-Agent as an observation to verify in live logs rather than as proof of identity.
Role and policy boundary
Atlassian's custom website connector is designed to crawl and index a website for use inside Atlassian's products. The documented outcome is not generic web search visibility: it is the ability to make selected website content available to Atlassian Search, Rovo Chat, and Rovo Agents. That makes the decision a content-governance question as much as a crawler question. A public documentation site may benefit from being discoverable in a customer's Rovo workspace, while a staging site, private support portal, or site containing customer-specific data should not be exposed through an unauthenticated crawl.
The public documentation gives important operational context but does not establish that every request carrying atlassian-bot is genuine. A User-Agent is self-declared and can be copied. If the request matters to an access-control decision, combine the header with source-IP verification from Atlassian's current documentation, normal edge logging, and application authorization. Never make a private page public merely because the crawler can reach it.
A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting. If you want the connector to read public documentation while excluding internal paths, use a dedicated group:
User-agent: atlassian-bot
Allow: /docs/
Disallow: /staging/
Disallow: /internal/
Disallow: /customer-data/
To request that the crawler not access the site, use:
User-agent: atlassian-bot
Disallow: /
Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect search, assistant, and training crawlers together, and it makes later audits harder because the owner of the decision is less specific.
Layered verification
Verify the same URL through each control plane instead of treating one successful fetch as authorization. Atlassian's connector documentation indicates that the connector expects a standard, top-level robots.txt at https://example.com/robots.txt; a robots file hidden behind a path, redirect chain, login, or application-generated challenge is not an equivalent control. Test the exact host and scheme that the connector will crawl.
The correct operational sequence is to inspect the crawler's request in edge logs, match the declared User-Agent, check the dedicated robots group and path, then inspect the returned HTML and headers. A permitted robots rule means that the site has published a crawl preference; it does not authenticate the caller. Conversely, noindex and X-Robots-Tag address search indexing semantics and should not be treated as a substitute for blocking access to confidential data.
The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why a path is allowed by robots.txt while a response header or application authorization layer still prevents useful indexing, rather than returning one blended website score.
For pages that should not be indexed or reused in search-oriented workflows, review the page-level directives as a separate signal:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives do not provide access control. If a response contains them, record the result as a policy conflict when the business intent is to keep the content out of a downstream index, and confirm how the consuming Atlassian feature interprets them before relying on that behavior.
WAF and Nginx remediation examples
To enforce a network-layer block for the declared crawler, match a normalized token rather than the entire browser-like User-Agent. This example is intentionally only a starting point: it blocks the self-declared string but does not prove that a request is Atlassian traffic.
{
"description": "Block declared Atlassian Rovo crawler",
"expression": "lower(http.user_agent) contains \"atlassian-bot\"",
"action": "block"
}
In Nginx, a local deny rule can be useful for a narrow path policy:
map $http_user_agent $deny_atlassian_bot {
default 0;
~*atlassian-bot 1;
}
server {
location /internal/ {
if ($deny_atlassian_bot) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not use a User-Agent match as the only protection for private content, and do not copy an if rule into a complex Nginx configuration without testing it in staging. If the objective is to keep public documentation available to Rovo while protecting sensitive areas, prefer a path-specific robots policy plus real authentication and WAF rate limiting.
Review checklist
After changing the policy, request the top-level /robots.txt directly and verify that the response is publicly reachable, served with a successful status, and contains the intended atlassian-bot group. Test one allowed documentation URL, one disallowed internal URL, and one redirect or canonical URL so the effective path is clear.
Then compare edge logs against the declared User-Agent, confirm that no confidential response is accessible without authentication, and inspect HTML metadata and X-Robots-Tag independently. Record the date, host, path, HTTP status, redirect chain, and source-IP evidence. Re-test after CDN or WAF changes, because a challenge page or a host-level redirect can make a previously valid connector configuration fail.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.