← Bot Directory/stepstoneCrawlBot
Bot directory / job-search

stepstoneCrawlBot: Robots.txt & Crawl Policy Reference

Technical reference for stepstoneCrawlBot, the job search crawler operated by The Stepstone Group, including its User-Agent, robots compliance, and job-indexing purpose.

AI Summary: stepstoneCrawlBot is the official crawler operated by The Stepstone Group to discover and index publicly accessible job postings. Suitable postings are published across its international platforms (like stepstone.de and totaljobs.com). The official documentation publishes the exact User-Agent and states that it strictly adheres to robots.txt directives and politeness rules. Because the operator does not publish a static IP list, use reverse/forward DNS verification before trusting the header.

Role and policy boundary

The Stepstone Group employs an automated web crawler to identify and analyze publicly accessible job postings on company websites and career portals. When a suitable job posting is detected, the crawler extracts information such as the job title, location, and description. These postings are then published free of charge on international platforms operated by The Stepstone Group (including stepstone.de, stepstone.at, stepstone.be, totaljobs.com, and irishjobs.ie).

The official documentation publishes this exact User-Agent string:

configuration / code
Mozilla/5.0 (compatible; stepstoneCrawlBot; +https://www.thestepstonegroup.com/english/crawler/)

The operator explicitly states that the crawler observes robots.txt directives and maintains appropriate intervals between requests to avoid placing undue load on servers. For selective access to public career pages:

configuration / code
User-agent: stepstoneCrawlBot
Allow: /careers/
Allow: /jobs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

For a complete exclusion from job indexing:

configuration / code
User-agent: stepstoneCrawlBot
Disallow: /

These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. The job-indexing purpose does not establish permission for third-party model training or generative AI input.

Layered verification

Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs. The reviewed Stepstone documentation does not publish a static IP list or a specific reverse-DNS verification procedure.

Because a User-Agent is easily spoofed, do not treat the header alone as authentication. Use standard network verification: perform a reverse DNS lookup on the source IP and look for a Stepstone-affiliated hostname, then perform a forward lookup to confirm it resolves back to the same IP. A stepstoneCrawlBot header from an unverified IP or a residential network should not be treated as authentic Stepstone traffic.

Compare observed traffic with the documented job-indexing role. Public HTML, career pages, metadata, feeds, sitemaps, and ordinary assets are consistent with discovery. Private endpoints, account routes, APIs, licensed content, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not permission. The operator states that the crawler maintains appropriate intervals between requests.

Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact stepstoneCrawlBot group you intend to publish. Check precedence, path matching, and actual request behavior. Meta directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

Begin in report-only mode and correlate the stepstoneCrawlBot User-Agent with reverse/forward DNS verification. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe stepstoneCrawlBot candidates before enforcement",
  "expression": "lower(http.user_agent) contains \"stepstonecrawlbot\"",
  "action": "log"
}

For a deliberately restricted private route, use the exact token only after verifying the source network:

configuration / code
map $http_user_agent $block_stepstone_private {
    default 0;
    ~*stepstoneCrawlBot 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_stepstone_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

This rule is route-scoped and does not prove identity. Do not invent an IP allowlist, Stepstone ASN exception, rate, or training permission without operator documentation.

A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action. The operator provides webcrawler@stepstone.de for questions regarding the crawler.

Test public career pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact stepstoneCrawlBot string, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Perform reverse and forward DNS verification to confirm a Stepstone hostname before trusting the client. Treat requests from unverified networks as spoofed.

Check your robots file for an exact stepstoneCrawlBot group. Test precedence and path behavior, and verify actual logs after publishing. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.

Review job indexing, AI input, reference use, and model-training decisions separately. The operator explicitly states the crawler is used to identify and publish job postings on Stepstone platforms.

Keep the profile at documented-limit: the operator publishes the exact User-Agent, robots compliance, and job-indexing purpose, but lacks a static IP list for verification on the reviewed page. Re-check the crawler documentation when the source network or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.

References

  1. Stepstone Crawler Information — official page documenting the User-Agent, robots.txt compliance, job-indexing purpose, and contact email; reviewed with the headless browser on 2026-08-25.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.