← Bot Directory/MatchorySearch
Bot directory / search-engine

MatchorySearch: Robots.txt & Crawl Policy Reference

Technical reference for the MatchorySearch registry label, including Matchory’s observed content signals and explicit limits around crawler identity and AI-use interpretation.

AI Summary: Matchory operates a supplier-data and procurement platform, but its current robots.txt does not list MatchorySearch as a named crawler. The file publishes Content-Signal: search=yes,ai-train=no,use=reference for all agents and separately blocks several named AI and commercial crawlers. Treat the registry User-Agent as unverified, interpret content signals narrowly, and do not assume search=yes permits AI-generated summaries or authenticates a requester.

Role and policy boundary

The registry describes MatchorySearch as a search-engine and content-discovery indexing crawler and records this User-Agent:

configuration / code
Mozilla/5.0 (compatible; MatchorySearch/1.3; +https://www.matchory.com)

Matchory’s official homepage is a live supplier-data layer for procurement. It describes resolved and enriched supplier profiles, global market knowledge, search and discovery, risk and compliance, and access through platforms or AI agents. This establishes the operator’s product context, but it does not independently confirm that the registered MatchorySearch token is a current production crawler.

The current official robots.txt publishes a Cloudflare-managed global signal:

configuration / code
User-agent: *
Content-Signal: search=yes,ai-train=no,use=reference
Allow: /

The comments define search as building a search index and providing hyperlinks and short excerpts, explicitly excluding AI-generated search summaries. They define ai-input as sending content into AI models for retrieval or generative search, ai-train as training or fine-tuning models, and use as immediate, reference, or full consumption. The file says restrictions expressed through these signals reserve rights under Article 4 of the European Union Digital Single Market copyright directive.

The same robots file separately disallows Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, CloudflareBrowserRenderingCrawler, Google-Extended, GPTBot, and meta-externalagent. It does not name MatchorySearch. Do not infer that the global signal authenticates a request or that search=yes grants permission for AI-generated summaries, model input, training, full-content use, or unrestricted reuse.

If logs confirm the exact token and your policy is to exclude it, a narrow defensive group could be:

configuration / code
User-agent: MatchorySearch
Disallow: /

For selective public discovery while protecting supplier intelligence and private accounts:

configuration / code
User-agent: MatchorySearch
Allow: /public/
Allow: /docs/
Disallow: /supplier-data/
Disallow: /account/
Disallow: /private/
Disallow: /api/

These are site-owner examples, not recovered MatchorySearch instructions. Robots.txt and Content-Signal are policy signals, not private-data access controls; use authentication, authorization, signed URLs, data minimization, and origin controls for sensitive supplier information.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Matchory’s reviewed homepage and robots file do not publish a MatchorySearch-specific IP range, reverse-DNS procedure, Crawl-delay, or crawler contact. The registry token is useful for triage but is not authentication.

Classify the requested use separately from the header. A public page and short search excerpt may fit the robots file’s narrow search=yes meaning. AI retrieval, grounding, generative summarization, model input, or training are distinct uses; the file does not grant them merely because search is allowed. The named ai-train=no signal should be treated as an explicit restriction for training/fine-tuning, while use=reference should not be extended to full-content use.

Compare observed behavior with a supplier-data discovery hypothesis without turning it into attribution. Public product, supplier, compliance, and market pages may be intended for search visibility. Private supplier records, contracts, internal notes, account routes, APIs, and bulk downloads require application authorization. High concurrency, repeated retries, or traffic that ignores signals may indicate spoofing, an integration, abuse, or another client. These observations establish impact and data risk, not operator identity or downstream use.

Evaluate /robots.txt independently. Confirm it is served from the correct host, returns a successful status and text content type, contains the exact groups you intend to publish, and is not overwritten unexpectedly by a CDN or managed security product. Test Content-Signal handling separately from Allow/Disallow; the comments define the semantic categories, but they do not authenticate the caller. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not secure private supplier intelligence or prove downstream compliance. Enforce sensitive boundaries in the application and at the origin, and document data-use permissions contractually where necessary.

WAF and Nginx remediation examples

If logs show the registry token, start with report-only mode and record the requested content and apparent use. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Review MatchorySearch content-use candidates",
  "expression": "lower(http.user_agent) contains \"matchorysearch\"",
  "action": "log"
}

After reviewing false positives and confirming the business decision, scope enforcement to supplier data and private routes:

configuration / code
map $http_user_agent $block_matchory_private {
    default 0;
    ~*MatchorySearch/1\.3 1;
}

server {
    location ~ ^/(supplier-data|account|private|internal|api)/ {
        if ($block_matchory_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and must not be the sole control for valuable procurement data. Do not invent an IP allowlist, reverse-DNS suffix, or managed verified-bot exception. If you want to block AI training or AI input, use the documented Content-Signal policy and provider controls in addition to route-level authorization; do not assume a robots directive alone removes previously collected data.

Test public supplier pages, search excerpts, structured data, sitemaps, private supplier records, accounts, APIs, exports, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, data-loss monitoring, contractual restrictions, and anomaly detection.

Review checklist

Search logs for the complete registered header and preserve source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the source can be independently verified and whether the observed use is search indexing, AI input, reference use, full-content use, or training; do not infer the use from the token alone.

Review the current Content-Signal definitions. search=yes is described as search indexing with hyperlinks and short excerpts, not AI-generated search summaries. ai-train=no explicitly restricts training/fine-tuning, while use=reference must not be treated as permission for full-content use. No MatchorySearch-specific group is currently documented in the reviewed file.

Protect supplier records, internal notes, contracts, accounts, and APIs with application authorization. Re-check Matchory’s homepage, robots.txt, and product documentation for a current bot identity, IP ranges, rate guidance, or support process. Keep this profile at documented-limit until the registry token is confirmed by Matchory.

Do not claim that a robots or Content-Signal change blocked a crawler or stopped downstream use from configuration alone. Verify subsequent logs and, where relevant, obtain a written operator response or contractual confirmation.

References

  1. Matchory homepage — official current supplier-data/procurement platform page, including supplier resolution, market knowledge, search/discovery, risk/compliance, and AI-agent access context; reviewed with the headless browser on 2026-08-25.
  2. Matchory robots.txt — official current file publishing Content-Signal: search=yes,ai-train=no,use=reference, global allow, and named-agent restrictions; reviewed with the headless browser on 2026-08-25.
  3. Robots Exclusion Protocol — standard for robots directives; it does not authenticate a User-Agent.
  4. Directive 2019/790, Article 4 — legal context referenced by Matchory’s Content-Signal comments; applicability depends on the operator’s circumstances and is not a substitute for legal advice.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.