Bot directory / search-engine

IndeedBot: Robots.txt & Crawl Policy Reference

Technical reference for the IndeedBot registry label, including the observed Indeed robots policy, source limitations, and safe verification controls.

AI Summary: IndeedBot 1.1 is a community-registry User-Agent label associated with Indeed job-search crawling, but the exact header was not independently confirmed in a current Indeed-owned bot page. Indeed’s live robots.txt was accessible and contains extensive global and named-agent rules, while the homepage was captcha-blocked during review. Treat the registry header as a log-search signal and verify current traffic before applying controls.

Role and policy boundary

The registry describes IndeedBot as an Indeed job-search web crawler and records this User-Agent:

configuration / code
Mozilla/5.0 (Windows NT 6.1; rv:38.0) Gecko/20100101 Firefox/38.0 (IndeedBot 1.1)

The registry source is a community-maintained crawler list, not an Indeed-owned policy page. A direct request to the Indeed homepage was blocked by a captcha during review, so the site’s general product behavior and bot documentation were not inferred from that page. The direct robots endpoint was accessible and contained a global group with Allow: /, numerous path-specific disallows, and multiple named-agent sections. The captured extract did not establish an exact IndeedBot-specific rule that can be safely attributed to this label.

Treat job-search indexing as the registry’s role hypothesis only. A matching header may come from a legacy client, a partner integration, a test harness, or a spoofed request. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private job-seeker data from the token.

If logs confirm this exact token and your site policy is to exclude it from discovery, a narrow defensive group could be:

configuration / code
User-agent: IndeedBot
Disallow: /

For selective access to public job listings while protecting accounts and application flows:

configuration / code
User-agent: IndeedBot
Allow: /jobs/public/
Allow: /job-postings/
Disallow: /admin/
Disallow: /account/
Disallow: /resume/
Disallow: /applications/
Disallow: /api/

These are site-owner examples, not recovered Indeed instructions. Robots.txt is advisory and cannot protect private resumes, applications, or licensed listings; use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No current Indeed-owned page reviewed in this pass published a bot-specific IP range, reverse-DNS procedure, or rate policy. The community registry header is therefore useful for triage but is not authentication.

Compare observed behavior with a job-search indexing hypothesis without turning it into attribution. Public job-listing HTML, canonical metadata, feeds, and ordinary assets may be consistent with indexing. Access to resumes, accounts, application endpoints, recruiter controls, or private APIs, as well as high concurrency, repeated retries, or traffic that ignores your restrictions, may indicate spoofing, abuse, a partner integration, or another client. These observations establish impact, not operator identity or downstream use.

Evaluate /robots.txt independently. Confirm the response is served from the correct Indeed-owned or target host, returns a successful status and text content type, and contains the exact IndeedBot group you intentionally publish. Do not assume that a global rule, a different named-agent group, or a path-specific rule automatically applies to this historical token. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate IndeedBot or secure private applicant data. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

If logs show the registry token, begin with a report-only rule and correlate it with source evidence. Adapt the expression to your WAF provider; it identifies an observed header only:

configuration / code
{
  "description": "Review observed IndeedBot candidate traffic",
  "expression": "lower(http.user_agent) contains \"indeedbot\"",
  "action": "log"
}

After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:

configuration / code
map $http_user_agent $block_indeed_private {
    default 0;
    ~*IndeedBot 1;
}

server {
    location ~ ^/(admin|account|resume|applications|private|internal|api)/ {
        if ($block_indeed_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and may catch a legitimate Indeed partner or a local test client. Do not invent an IP allowlist or reverse-DNS suffix from the community registry or the current robots file. Test public job listings, feeds, sitemaps, media, resumes, applications, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search logs for the complete registry header and the simpler IndeedBot token. Preserve representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether traffic stays within public listing routes and whether its source can be independently verified; do not classify it as authentic from the header alone.

Re-check the live robots.txt for a current exact IndeedBot group before relying on any path rule, because a large file may contain several unrelated named-agent sections. Treat the homepage captcha as an access limitation, not evidence of bot activity or inactivity. Keep this profile at partially-documented until Indeed publishes a current canonical bot page, exact User-Agent, source-verification method, IP ranges, rate guidance, or explicit opt-out instructions.

Decide whether your objective is to preserve job-search visibility, limit listing extraction, protect applicant data, or reduce crawl load. Publish exact robots rules and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent access logs and response behavior.

References

  1. Indeed robots.txt — official current robots file observed with a global group, extensive path controls, and multiple named-agent groups; no exact IndeedBot-specific rule was established in the captured review on 2026-08-25.
  2. Indeed homepage — direct headless-browser request was captcha-blocked during review; no general page claims were inferred.
  3. Crawler User Agents community registry — community source for the registry User-Agent; it does not authenticate Indeed infrastructure.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.