semantic-visions: Robots.txt & Crawl Policy Reference
Technical reference for the semantic-visions-discovery User-Agent associated with Semantic Visions' svEye open-source intelligence platform. Learn how to verify the token safely.
AI Summary: Semantic Visions' official site documents svEye as a platform that gathers and enriches online news and blog data for real-time market intelligence. It does not publish a crawler contract for the registry token
semantic-visions-discovery, so the active identity, source networks, rate, robots behavior, and training use are unverified. Treat the token as a log observation and enforce unwanted traffic with layered controls.
Role and policy boundary
Semantic Visions' official product page says svEye automatically identifies, classifies, and gathers online news and blogs across millions of sources and multiple languages. It describes real-time monitoring, contextual enrichment, historical screening, and predictive insights for organizations analyzing companies, commodities, industries, locations, and events. This establishes a documented web-monitoring product context.
It does not establish that every request declaring semantic-visions-discovery is operated by Semantic Visions, nor that the token is currently active. The public product page reviewed for this profile does not publish a dedicated User-Agent contract, source IP ranges, crawl rate, robots behavior, or webmaster contact. Do not convert the product description into a claim that a particular request is collecting foundation-model training data.
If your logs confirm the exact token and you want to communicate a restriction, publish:
User-agent: semantic-visions-discovery
Disallow: /
For a selective policy that leaves a public newsroom available but excludes private or high-cost routes:
User-agent: semantic-visions-discovery
Allow: /news/
Allow: /public/
Disallow: /private/
Disallow: /drafts/
Disallow: /internal/
Disallow: /api/
Robots.txt is advisory. Protect private articles, customer data, paywalls, and APIs with authentication and authorization.
Layered verification
Start with raw access logs and preserve the full User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirect chain, timestamp, and request rate. The registry header is self-declared; no official IP list or source-verification method was found.
Analyze behavior without over-assigning purpose. A market-intelligence service may request news pages, feeds, sitemaps, and archives, while an abusive scraper may generate high concurrency, repeated retries, or requests for private and API routes. These observations demonstrate operational impact but cannot prove the operator or downstream use.
Evaluate /robots.txt independently. Confirm the canonical host, status, content type, exact group, and path match. For page-level discoverability, inspect:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives communicate indexing preferences but do not protect confidential content. Use application authorization, signed URLs, and origin controls for sensitive data.
WAF and Nginx remediation examples
If logs confirm unwanted requests carrying the token, a narrow WAF rule can block the declared identity:
{
"description": "Block observed Semantic Visions discovery token",
"expression": "lower(http.user_agent) contains \"semantic-visions-discovery\"",
"action": "block"
}
For Nginx, scope enforcement to private or high-cost paths while investigating public news access:
map $http_user_agent $block_semantic_visions {
default 0;
~*semantic-visions-discovery 1;
}
server {
location ~ ^/(private|drafts|internal|paywall|api)/ {
if ($block_semantic_visions) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof or evade and may block an approved monitoring integration. Test ordinary browsers, feed readers, social previews, search crawlers, and newsroom tools. Pair edge matching with rate limits, authentication, paywall enforcement, and anomaly detection rather than treating the header as identity proof.
Review checklist
Search logs for the exact semantic-visions-discovery value and record paths, volume, response sizes, source networks, status, and timing. Compare observations with Semantic Visions' current svEye product description, but keep product capabilities separate from crawler attribution. No official source range or dedicated crawler policy was found during this review.
Decide whether your objective is to preserve news and market-intelligence visibility, prevent possible content collection, protect embargoed material, or reduce crawl load. Publish a targeted robots group for the exact observed token, enforce private and paywalled routes with WAF and application controls, and test articles, feeds, sitemaps, media, archives, and APIs separately. Revisit the profile if Semantic Visions publishes a verifiable crawler User-Agent, source method, rate policy, or robots statement.
References
- Semantic Visions — official svEye product page documenting automated online-news and blog collection and enrichment.
- Semantic Visions crawler URL — no dedicated crawler contract was published on the official page reviewed on 2026-08-24.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.