Jooblebot: Robots.txt & Crawl Policy Reference
Technical reference for Jooblebot, including its published job-indexing role, exact User-Agent, observed robots.txt boundaries, and limits around source verification.
AI Summary: Jooble documents Jooblebot as the crawler that indexes jobs from the web and sends users to the original listing. The official bot page publishes the core User-Agent, while the current robots.txt contains extensive global and named-agent rules but no Jooblebot-specific group, Crawl-delay, IP range, or reverse-DNS procedure in the reviewed content. Treat the header as a useful signal, not authentication, and make robots decisions from an exact current rule.
Role and policy boundary
Jooble’s official bot page says Jooble indexes jobs from the web using a web crawler that identifies itself as JoobleBot or the following core string:
Mozilla/5.0 (compatible; Jooblebot/2.0; +http://jooble.org/jooble-bot)
The inventory includes a longer browser-compatible form:
Mozilla/5.0 (compatible; Jooblebot/2.0; Windows NT 6.1; WOW64; +http://jooble.org/jooble-bot) AppleWebKit/537.36 (KHTML, like Gecko) Safari/537.36
Jooble says a short extract from a job is displayed in its results and that selecting the job takes the user to the full listing on the originating site. This establishes a job-search indexing role. It does not grant access to private applications, candidate profiles, employer accounts, recruiter APIs, or licensed data outside the public listing boundary.
The current Jooble robots file contains a global group beginning with User-agent: * and lists many account, employer, redirect, tracking, search-result, and subscription paths as disallowed. It also names several specific agents in separate groups. The reviewed file did not contain a Jooblebot-specific group, Crawl-delay, source IP range, or reverse-DNS verification procedure. Do not assume that another named-agent group or the global group expresses a complete Jooblebot policy without checking the parser’s exact matching behavior.
If logs confirm the core Jooblebot token and your site policy is to exclude it, a narrow defensive rule could be:
User-agent: Jooblebot
Disallow: /
For selective access to public listings while protecting applications and accounts:
User-agent: Jooblebot
Allow: /jobs/public/
Allow: /job-postings/
Disallow: /admin/
Disallow: /account/
Disallow: /resume/
Disallow: /applications/
Disallow: /api/
These are site-owner examples, not a complete Jooble instruction. Robots.txt is advisory and cannot protect private resumes, applications, or licensed listings; use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Jooble’s reviewed bot page and robots file publish no IP range or reverse-DNS verification method. The header is therefore useful for triage but is not authentication.
Compare observed behavior with the documented job-indexing role. Public job-listing HTML, canonical metadata, feeds, and ordinary assets may be consistent with indexing. Access to resumes, accounts, applications, employer controls, or private APIs, as well as high concurrency, repeated retries, unexpected downloads, or traffic that ignores your restrictions, may indicate spoofing, abuse, a partner integration, or another client. These observations establish impact, not operator identity or downstream use.
Evaluate /robots.txt independently. Confirm that the response is served by the intended host, returns a successful status and text content type, and contains the exact Jooblebot group you intentionally publish. Test the full browser-compatible header, the core Jooblebot token, and the global group separately; a crawler may match the first applicable rule under the parser’s semantics. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate Jooblebot or secure private applicant data. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If logs show Jooblebot traffic, begin with a narrow report-only rule and correlate it with source evidence. Adapt the expression to your WAF provider; it identifies an observed header only:
{
"description": "Observe Jooblebot job-indexing candidates",
"expression": "lower(http.user_agent) contains \"jooblebot\"",
"action": "log"
}
After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:
map $http_user_agent $block_jooble_private {
default 0;
~*Jooblebot/2\.0 1;
}
server {
location ~ ^/(admin|account|resume|applications|private|internal|api)/ {
if ($block_jooble_private) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof and may catch a legitimate Jooble partner or a local test client. Do not invent an IP allowlist or reverse-DNS suffix from the current pages. Test public listings, feeds, sitemaps, media, resumes, applications, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for both the complete registry header and the core Jooblebot/2.0 token. Preserve representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the traffic stays within public job-discovery routes and whether its source can be independently verified; do not classify it as authentic from the header alone.
Re-check the official bot page and robots.txt for a Jooblebot-specific group before relying on any global or unrelated named-agent rule. Confirm that a selected robots group protects accounts, applications, resumes, employer APIs, and tracking endpoints without unintentionally removing public job visibility. Keep IP and DNS verification marked unavailable until Jooble publishes a current method.
If you need an inclusion or indexing decision, use Jooble’s published contact path rather than assuming that robots changes are acknowledged immediately. Re-check the bot page for changes to the User-Agent, result behavior, and contact process. Do not claim successful blocking from configuration alone; verify subsequent access logs and response behavior.
The documented purpose is job indexing and referral to source listings. It does not establish AI-model training use; keep that boundary separate from any future service or client using a similar header.
References
- What is JoobleBot? — official page documenting Jooble’s job-indexing role, core User-Agent, result extracts, source-link behavior, and contact path; reviewed with the headless browser on 2026-08-25.
- Jooble robots.txt — official current file with global and named-agent rules, observed with the headless browser on 2026-08-25; no Jooblebot-specific group, IP range, or DNS verification method was established.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.