7Maps · Methodology

How 7Maps checks MCP servers

Every request, interval, timeout, rule and formula behind the map, taken from the code that runs it. If this page and the data ever disagree, the page is wrong and gets a new version.

Methodology v1.1 (2026-10-03)JSON twinChangelog

What is checked

Every remote server in the official MCP registry (registry.modelcontextprotocol.io/v0.1/servers, latest version of each entry, entries marked deleted left out), plus servers their owners submitted that passed a week of observation. For each server the registry entry's first remote of type streamable-http is used, or else its first remote.

7Maps-bot sends what any MCP client sends when it connects, and nothing more:

  1. initialize, offering protocol version 2025-06-18, as client 7maps-bot 0.2.
  2. notifications/initialized.
  3. tools/list, following nextCursor for up to 3 pages. At most 500 tools are kept per server.

Once a day per origin it also reads /robots.txt, /.well-known/oauth-protected-resource (does sign-in publish OAuth metadata) and /.well-known/agent-card.json or /.well-known/agent.json. It never calls a tool, never signs in and sends no credentials. User agent: 7Maps-bot/0.2 (+https://7it.co.il/7maps/bot/).

Left out before any request: remotes that are SSE only, addresses with a template ({ or }), anything not https, IP addresses, localhost, .local and .internal names, hosts that opted out, and origins whose robots.txt says Disallow: / to 7Maps-bot or to *.

A redirect on initialize to the same registrable domain is followed once. A redirect to another domain is recorded, not followed. For servers with a high-risk tool, and for submitted servers, the age of the domain is read from its public RDAP record (rdap.org, at most 40 new lookups per scan, kept 30 days); that lookup never reaches the server.

From where, and how often

  • Daily, every server. A job runs at minute 5 of every UTC hour and checks one of 24 shards (shard = FNV-1a hash of the server's host and path, mod 24). The shard for an hour moves by one each day (shard = (UTC hour + days since 2026-10-03) mod 24), so each server is checked once a day, one hour earlier than the day before, and after 24 days has a result at every hour of the day. 50 checks run at a time.
  • Hourly, the most used servers. After each shard, up to 300 servers are checked again: those agents reported on in the last 7 days first, then those listed in the Claude or ChatGPT directories. Only servers a daily check reached in the last 2 days, never opted-out hosts or known gateways. 25 at a time. A result older than 2 hours is ignored.
  • Every 5 minutes, servers whose owner verified control. Up to 3 per owner by default (an owner is the GitHub user for io.github.<user>/... names, otherwise the registrable domain, or the full host on shared hosting such as vercel.app, workers.dev or onrender.com), earliest verified first, 20 at a time within a 50-second run, each check capped at 30 seconds. A server not reached in time is skipped, not counted as down.

All checks run from one cloud region in the Washington, D.C. area (Vercel iad1), so times include the distance from there.

Also: servers submitted by their owners are checked every 6 hours for 7 days and join the map with at least 20 checks, 90% answered with tools, at most one tool-list change and every tool described (once, 7 more days when 70% to 90% answered). npm and PyPI packages that servers ship as are read once a day from public package metadata, and their current versions are checked against OSV at 04:05 UTC. Compute caps: $0.10 a day for the daily checks (past it, the remaining shards wait for the next day, and a skipped day is not recorded as down) and $0.25 a day by default for the five-minute checks.

Timeouts

robots.txt
3 seconds (no answer means allowed)
initialize
8 seconds
notifications/initialized
5 seconds
tools/list
8 seconds per page, up to 3 pages
well-known files per origin
4 seconds each
five-minute check
30 seconds overall, on top of the limits above

What counts as answering, sign-in needed, or down

Answered with tools o
initialize returned a result and tools/list returned a result (an empty list included).
Asked for sign-in a
HTTP 401 or 403 on initialize, or on tools/list before any tool was listed.
Asked for payment p
HTTP 402 on initialize. Where the server asks agents to pay is recorded; a change of that address is reported.
Not answering d
A timeout, no connection, a redirect to another domain, any other HTTP error, a JSON-RPC error on initialize, an answer that is not MCP, or no tool list after a successful initialize.

Sign-in and payment answers count as up: the server is there and answering. The handshake time is the wall time of the whole check (initialize, notifications/initialized and tools/list).

On the server page, badge and widget, an HTTP 400, 405, 406, 415 or 422, or an answer that is not MCP, is shown as "answers, but not as a standard MCP endpoint", never as not answering or down (the daily history still records it as d). A failed five-minute or hourly check changes the shown status only from the second failure in a row, the same rule as owner alerts; a single failure falls back to the check before it.

How tool risk levels are assigned

read-only an agent may use it alone. changes data needs a person's approval. high risk deletes, pays, grants, deploys or runs code.

The level comes from the tool's own annotations, the verb in its name and its description. The verb is the first word of the name, or the last when the first is not a known verb (delete_page, pages_create); words split on _ - . / : and camelCase. The first rule that applies wins:

  1. destructiveHint: true (and not readOnlyHint: true): high risk.
  2. The verb is a high-risk verb: high risk.
  3. readOnlyHint: true: read-only.
  4. The name holds a sensitive noun (payment, payout, refund, transfer, password, credential, secret, API key, shell, sudo, terminal, wallet) and the verb is not a reading verb: high risk.
  5. The verb is a changing verb: changes data.
  6. The verb is a reading verb, the tool has annotations and does not say readOnlyHint: false: read-only.
  7. The verb is a reading verb but the server publishes no annotations: changes data, until someone checks it.
  8. readOnlyHint: false, or nothing clear: changes data.

Then, unless the server marked the tool read-only, a description that plainly says it moves money, charges a card, deletes permanently or deletes the, all or a record, drops a table or database, executes arbitrary, shell, SQL or code, runs shell or arbitrary commands, grants access or permissions, or changes passwords raises it to high risk.

The verb lists

High risk: delete, remove, destroy, drop, truncate, wipe, purge, erase, revoke, grant, transfer, refund, charge, pay, payout, purchase, buy, checkout, order, deploy, publish, release, merge, execute, exec, run, sign, ban, suspend, terminate, cancel, close, rotate, reset, kill, shutdown.

Changes data: send, email, mail, message, sms, notify, post, reply, comment, tweet, create, update, edit, write, upload, schedule, book, invite, assign, share, move, rename, set, add, insert, modify, patch, commit, push, submit, approve, reject, label, tag, archive, import, sync, enable, disable, install, save, store, put, upsert, append, replace, restore.

Reading: get, list, read, search, fetch, find, query, describe, view, show, check, analyze, analyse, summarize, summarise, lookup, retrieve, count, status, preview, estimate, validate, inspect, browse, explain, compare, test, ask, resolve, download, export, calculate, map, scan, lint, diff.

The exact patterns are in the JSON twin.

This is a heuristic, read from what a tool says about itself. It has known false positives (a tool named run_report only reads, yet its verb makes it high risk) and false negatives (a tool with a mild name that does something risky and does not say so). It is not a security audit, and a level is never a judgement of a server or of the people who run it.

Findings shown per server

Heavy tool list
The tool list is over 40,000 characters (about 10,000 tokens).
Very long description
A tool description is over 2,000 characters.
Missing descriptions
Some tools have no description.
No annotations
No tool declares annotations.
Missing schema
Some tools declare no input schema.
Duplicate names
Two tools share a name.
Older protocol
The server answers with a protocol version older than 2025-03-26.
Slow
The whole check took over 3 seconds.

Two pattern checks on tool descriptions (instruction-like wording, invisible characters) are kept for a person to review and are never published or shown to agents: on 2026-10-02 every match sampled was a false positive.

How token cost is estimated

The size of a tool list is the length in characters of the JSON of the tools array exactly as tools/list returned it (all pages read). Tokens are estimated as that length divided by 4. Real counts depend on each model's tokenizer, so this is an estimate, used the same way for every server.

How changes are detected

Each tool gets a hash of its name, description, input schema and annotations; the tool list gets a hash of all of them. When a server's list hash differs from its previous daily check, the difference is recorded tool by tool: added, removed, changed, a tool that can change things rewritten at the same level, or a rise in risk (a tool added as high risk, or a level going up). A change while the server reports the same server version is marked silent. Also recorded: started or stopped answering, now asking for sign-in, a new payment address, a redirect to another domain. Change days are kept for 60 days per server.

Uptime

Each server keeps its last 30 daily results, one letter per calendar day (o, a, p or d as above). A day on which the server's daily check did not run (a skipped run, the daily cost cap) is n, no data: it keeps every other result on its own date and counts neither as up nor as down. Days observed are the days that are not n.

daily uptime %     = 100 x (days observed not d) / (days observed)
five-minute 30d %  = 100 x (answered checks) / (checks), last 30 UTC days
five-minute 90d %  = the same over 90 days, only when older checks exist

One decimal. A server that missed any check is never rounded up to 100% (it shows at most 99.9%). The five-minute figure is used once a server has at least 12 checks (an hour); otherwise the daily figure, labelled as daily. Hourly checks update the current status but are not counted in uptime. A figure is called a 30-day uptime only from 14 observed days; before that it is labelled "over N days observed". The uptime distribution on the state page appears once the median server has 14 observed days; until then the page gives the days observed and the date it appears.

Chance the next call works, and stability

success chance = (1 + sum of w over answering days) / (2 + sum of w over observed days)
                 w = 0.5 + i / n for day i of n, oldest i = 1, newest weighs most (n days count in n, add nothing)
stability      = 1 / (1 + 0.5 x days the tool list changed in the last 30)

Both measure the handshake and tool list, not a tool call. The hour-of-day profile keeps the latest result at each UTC hour and is shown once 6 hours have one.

Most reliable by category

A server is listed when it is answering with tools right now (the hourly check when newer, else the daily check), has a category read from its tools, and has at least 14 observed daily results (of up to 30). Order:

  1. Share of observed days (not n) that are not d, highest first.
  2. Newest handshake time, lowest first.
  3. More days observed.
  4. The server's address, so ties are stable.

Top 10%: a place within the first tenth of the ranked servers, rounded up, in categories with at least 10 ranked servers. Nothing else changes the order: placement is never sold, and plans, owner verification and agents' ratings do not count. The lists.

How route picks a server and tool

score = text match x success chance x stability x least privilege
        x (1 - 0.3 x share of the server's tools that are high risk)
        x verb bonus x directory factor x guard x agents factor
Text match
Words of the task against the tool name (counted double) and description (length-normalized), each word weighted by how rare it is across all tools.
Verb bonus
1.6 when the task's action verb is in the tool name.
Least privilege
0.35 when the tool can do more than the task needs, 0.15 when it can do less.
Directory factor
1.15 when the server is listed in the Claude or ChatGPT directory.
Guard
0.5 for a read-only task on a server that also has high-risk tools, 0.7 for gateways, 0.8 for servers submitted in the last 30 days or domains younger than 90 days.
Agents factor
0.5 + 0.5 x reliability once the server has stars (see ratings), otherwise 1.

Set aside and named as such: servers not answering with tools, with a tool-list change today, with a live incident, or with a rise in risk in the last 7 days when the task changes things. For high-risk tasks also new servers, young domains, gateways and servers that redirect to another domain.

Ratings from agents and gateway sensors

Agents report how a call went with report_road or the last_trip field on a paid call. Fixed fields only, nothing written in the agent's own words: worked or failed, one of seven failure reasons (unreachable, auth, args, server_error, timeout, rate_limited, wrong_result), milliseconds, whether the result matched the description, result tokens, amount charged, one of four surprises (side_effect, unexpected_cost, data_leak_suspected, other) and the tool-list hash the agent saw.

reporter weight = (paid ? 1.0 : 0.3) x (0.5 + 0.5 x (agree + 1) / (agree + disagree + 2))
decay           = 0.5 ^ (age / 7 days), reports from the last 30 days
reliability     = sum(weight x decay x ok) / sum(weight x decay)
stars           = 1 + 4 x (0.45 reliability + 0.2 accuracy + 0.2 (1 - surprise share) + 0.15 speed)
speed           = clamp(1 - (median ms - 500) / 5000, 0, 1), 0.7 when no times

A report agrees with our checks when its result matches whether our newest check saw the server answering, or when it reports a failure other than unreachable or timeout (a server_error only when our check saw a server error too); a disagreement counts half. Accuracy and surprise share come from the newest 30 reports. Limits: 50 reports a day per reporter and 3 per server.

A reporter counts as paid only with a license key that validates as active or a payment that settled; a report sent with a call that was not charged counts as anonymous. Anonymous reporters are told apart by network address (IPv4 /24, IPv6 /48), never by user agent. Stars, and the reliability share shown publicly, appear from 3 distinct reporters and a decayed weight of at least 1.0 from paying reporters; otherwise "not enough reports".

Gateway sensors (opt-in, open source) count like an anonymous agent: one entry per sensor, server and day, weight 0.3 times the sensor's agreement with our checks, contributing its share of calls that worked. A sensor counts as one reporter. Up to 20,000 outcomes per sensor per day by default, 100 per batch, outcomes older than 24 hours dropped.

When a live incident is shown

A failure counts when it was reported in the last 2 hours, after the last call anyone saw work. For a sensor, its newest outcome on the server failed.

An incident is shown only when it is confirmed, by one of: our own check in the same 2 hours saw a real failure (a timeout, no connection, an HTTP 5xx or a broken MCP handshake; a 404, a redirect, sign-in or another 4xx is not one), from the hourly check, the five-minute check or that day's daily check; or a paying reporter is among those reporting the failure. Confirmed, it still needs two signals: two distinct failing reporters, or one failing reporter and our own check.

Anonymous agents alone, sensors alone, or both together never open an incident. Failing sensors join a confirmed incident when our check confirms it (a sensor whose only reason is server_error, only when our check saw a server error too), or when a paying reporter gave the same reason (never server_error). Unconfirmed failures only lower the internal rating and are reviewed by a person. An incident ends when a call works again, when the 2 hours pass, or when a person closes it after review.

Signed daily history, and how to verify it

At 01:05 UTC the previous UTC day is sealed. Every file of that day (the index, each shard's results, tool catalogs, change logs, run summaries, the package layer and the day log of agent reports) becomes one line, sha256  path, sorted by path. Then:

root      = sha256( previous root (or "") + "\n" + lines joined with "\n" )
signature = Ed25519 over the 32 bytes of the root, base64

If the 01:05 run does not happen, a later hourly run that day signs it, and any of the 7 days before that has no root is signed then too, oldest first. A day that has a root is never signed again.

The previous root is the newest one in the 14 days before, so each day is chained to the one before it and history cannot be rewritten without breaking every later root. Roots: https://7it.co.il/api/maps?a=root&day=YYYY-MM-DD. Public key: integrity.public_key_pem in /.well-known/7maps.json.

// Node 18+: node verify.mjs
import { createHash, createPublicKey, verify } from 'node:crypto';
const day = '2026-10-02';
const wk = await (await fetch('https://7it.co.il/.well-known/7maps.json')).json();
const r = await (await fetch('https://7it.co.il/api/maps?a=root&day=' + day)).json();
const root = createHash('sha256').update((r.previous_root || '') + '\n' + r.files.join('\n')).digest('hex');
const ok = verify(null, Buffer.from(r.root, 'hex'), createPublicKey(wk.integrity.public_key_pem), Buffer.from(r.signature, 'base64'));
console.log(root === r.root && ok ? 'verified' : 'mismatch');

Checking a single file against its line needs that file; ask through the contact address below. Not in the roots: hourly and five-minute checks, sensor data, owner files and demand counts.

Data retention

What is never collected

To run the service and stop abuse, the 7Maps MCP server keeps a log of calls (tool, time, client name, country, and an agent id that is a keyed hash of address and user agent). It is not published. This site's pages use Google Analytics.

Opt-out and corrections

Opt out: add User-agent: 7Maps-bot and Disallow: / to the host's robots.txt (works at once), or use the form on the 7Maps-bot page, which asks for a short token at /.well-known/7maps-optout.txt so nobody can remove a server they do not control. The host is never contacted again and drops off the published map at its next daily check. Past dated files stay as they were, because signed history is not rewritten. Removal is never charged.

Corrections: owners can verify control to add a short note in their own words and get five-minute checks. Anyone can write through the contact address on 7it.co.il. A wrong rule is fixed in the code and recorded below.

Changelog

Any change to a check, interval, timeout, rule, formula, weight or retention gets a new version here and in the JSON.

v1.1 · 2026-10-03

Daily history: one code per calendar day, with n for a day without a daily check (it used to gain one code per check, so a skipped check moved older results off their dates); uptime, success chance and the reliable lists count observed days only. Signed roots: a missed day is signed by a later run (yesterday and the 7 days before), never re-signed. Retention: raw sensor, demand, view, ask, badge and daily-call files now expire on a schedule (see Data retention).

v1.0 · 2026-10-03

First published version.

7Maps by 7IT. Observations with a date, not a judgement. Also: State of MCP uptime · Most reliable · Compare · Tool search evaluation · 7Maps-bot.