Qwantbot: Robots.txt & Crawl Policy Reference
Technical reference for Qwantbot and Qwantbot-news, including Qwant’s published User-Agent pattern, robots behavior, DNS verification, and refreshable IP snapshot.
AI Summary: Qwant’s first-party crawler page documents Qwantbot and Qwantbot-news as search crawlers, publishes the pattern
Mozilla/5.0 (compatible; Qwantbot{-news}/X.Y_{worker_id}; +https://help.qwant.com/bot/), and says the robots standard is respected. Qwant recommends reverse DNS ending inqwant.com, optional forward confirmation, or its refreshable JSON IP list. The reviewed snapshot lists194.187.171.0/24; do not treat that prefix as permanent or use User-Agent alone as authentication.
Role and policy boundary
Qwant’s official help page says its web crawlers enhance the Qwant index and provide its search service. It identifies the stable marker Qwantbot and documents a family of User-Agents:
Mozilla/5.0 (compatible; Qwantbot{-news}/X.Y_{worker_id}; +https://help.qwant.com/bot/)
The brace-delimited -news and worker identifier portions are optional according to the page. Published examples include:
Mozilla/5.0 (compatible; Qwantbot/1.0_12345; +https://help.qwant.com/bot/)
Mozilla/5.0 (compatible; Qwantbot-news/2.0; +https://help.qwant.com/bot/)
The registry stores Mozilla/5.0 (compatible; Qwantbot/1.0_800166; +https://help.qwant.com/bot/). That value fits Qwant’s published family, but keep the registry version distinct from any current observed version and from Qwantbot-news. Do not assume every client containing the word Qwantbot is genuine.
Qwant states that its crawler respects the robots rules standard. On your own site, a specific exclusion can be expressed as:
User-agent: Qwantbot
Disallow: /private/
Disallow: /admin/
Disallow: /account/
Disallow: /api/
If you need to include the news variant, publish a separate group or verify your parser’s matching behavior:
User-agent: Qwantbot-news
Disallow: /private/
Disallow: /licensed/
A complete exclusion is also possible:
User-agent: Qwantbot
Disallow: /
These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. Qwant’s search-crawler documentation does not establish permission for AI input, reference use, or model training.
Layered verification
Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs. Qwant’s preferred verification method is reverse DNS: a crawler IP should resolve to a hostname ending with qwant.com. Optionally perform a forward lookup of that hostname and confirm that it resolves back to the same observed address.
Qwant’s page gives this illustrative Linux flow:
host 91.242.162.1
# 1.162.242.91.in-addr.arpa domain name pointer qwantbot-1-162-242-91.qwant.com.
host qwantbot-1-162-242-91.qwant.com
# qwantbot-1-162-242-91.qwant.com has address 91.242.162.1
Treat the example as methodology, not as a current allowlist. DNS evidence must be checked at observation time, and the forward result must include the original source IP. A User-Agent or a reverse-DNS suffix without forward confirmation is not sufficient authentication.
Qwant publishes an alternative JSON IP list at qwantbot.json and recommends refreshing it daily because it can change at any time. The reviewed snapshot had creationTimestampSecond: 1764143556 and 194.187.171.0/24 as its IPv4 prefix. This is a dated snapshot, not a permanent network claim. Store the retrieval time and compare the list again before applying a restrictive rule.
Compare observed traffic with the documented search-crawler role. Public HTML, metadata, links, feeds, sitemaps, and ordinary assets may be consistent with search indexing. Private endpoints, account routes, APIs, licensed content, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not permission for unrelated data uses. Qwant’s crawler page does not document model-training or downstream AI-use policy.
Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact Qwantbot group you intend to publish. Check precedence, path matching, and actual request behavior. Meta directives can express indexing preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
Begin in report-only mode and correlate the complete Qwant User-Agent with reverse/forward DNS and the current IP JSON snapshot. Adapt the expression to your WAF provider:
{
"description": "Observe Qwantbot candidates before enforcement",
"expression": "lower(http.user_agent) contains \"qwantbot\"",
"action": "log"
}
For a deliberately restricted private route, use the official family marker only after verifying the source network and current policy:
map $http_user_agent $block_qwant_private {
default 0;
~*Qwantbot(?:-news)?/ 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_qwant_private) { return 403; }
try_files $uri $uri/ =404;
}
}
This rule is route-scoped and does not prove identity. Do not hard-code the reviewed 194.187.171.0/24 prefix indefinitely; refresh Qwant’s JSON and validate DNS at request time or through a controlled allowlist process. Do not invent additional IP ranges, ASN exceptions, rates, or training permissions.
A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action. Contact qwantbot@qwant.com for crawler problems using representative logs and timestamps.
Test public pages, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, Qwantbot-news handling, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.
Review checklist
Search logs for Qwantbot and Qwantbot-news, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. Compare the observed header to Qwant’s published family pattern rather than only the registry’s 1.0_800166 example.
Perform reverse DNS and confirm a qwant.com suffix, then perform the optional forward lookup back to the same IP. Compare against the current official JSON list when DNS is unavailable or as a secondary check. The reviewed snapshot listed 194.187.171.0/24 with a creation timestamp; refresh daily because Qwant says the list can change at any time.
Check your robots file for exact Qwantbot and, if relevant, Qwantbot-news groups. Test precedence and path behavior, and verify actual logs after publishing. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.
Review search indexing, AI input, reference use, and model-training decisions separately. Qwant documents search crawling and robots behavior, but that does not establish downstream permission or use for an individual request.
Keep the profile at documented: Qwant publishes the crawler family, robots behavior, DNS verification method, a refreshable IP source, and a support contact. Re-check the page, JSON snapshot, and DNS evidence when the User-Agent or source network changes. Do not claim successful blocking or verification from configuration alone; validate later logs.
References
- Qwant Web crawler — first-party page documenting Qwantbot/Qwantbot-news patterns, robots behavior, reverse/forward DNS verification, refreshable IP JSON, and qwantbot@qwant.com; reviewed with the headless browser on 2026-08-25.
- Qwantbot IP JSON — official snapshot with
creationTimestampSecond: 1764143556and IPv4 prefix194.187.171.0/24, reviewed on 2026-08-25; Qwant recommends daily refresh. - Robots Exclusion Protocol — standard referenced by Qwant’s crawler page.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.