7Maps · Tool search evaluation

How well does 7Maps find the right MCP tool?

A fixed set of plain-language agent jobs, criteria written before any result was seen, and a judge with no model in it. The score is measured again every month and published here, good or bad.

October 2026 Evaluation set v1, run on the 1st of each month.

0.77P@1: the first result fits the job
0.68P@5: share of the top 5 that fit
0.69nDCG@5: order quality of the top 5
0.86MRR: how high the first fit is
98%jobs with a fitting tool on an answering server in the top 5
700 msmedian time to answer, measured from GitHub's US runner or our machine

Scores are between 0 and 1 (higher is better). Searched: every tool of every server on the map that answered with a tool list.

Month by month

Loading the history.

P@5nDCG@5

Each point is one monthly run of the same evaluation set. Data: history.json.

By kind of job

nDCG@5 for each category of the latest run.

How it is measured

What is searched
The public 7Maps tool search, /api/toolsearch, the same search the find_tool tool uses, with its defaults: any risk level, answering servers only, 10 results. One request per job, from a script, on production.
The jobs
Over 100 short jobs an agent might be given, in plain words, across email, calendar, databases, web search, payments, files, code, CRM, docs, maps and geo, weather, translation, scraping, images, analytics, ticketing, chat, storage, sign-in, monitoring and more. The set as JSON (CC BY 4.0).
What counts as a fit
Each job lists groups of synonyms. A result fits when its tool name plus description contains a word or phrase from every group (light stemming, so "sends" matches "send"), and when its risk level is within what the job allows: a read job accepts only read-only tools. Some jobs add a "prefer" group (for example "postgres" for a Postgres job): a fit that also matches it gets the higher grade.
No model in the judge
The judge is a short deterministic function in scripts/eval-toolsearch.mjs. Anyone can run the same jobs against the same search and get the same grades for the same results.
Metrics
P@1 and P@5 count fitting results in the top 1 and top 5. nDCG@5 rewards fitting results near the top, against an ideal list of five top-grade fits. MRR is 1 divided by the rank of the first fit (top 10). Latency is the median full request time.
Limits
Words are a proxy: a tool can match every word and still do the job badly, and a good tool described in unusual words can be missed. The risk level is the one 7Maps assigns. The score says how often the search puts a plausible tool on top, not whether a call would succeed.

The rule: no tuning to the test

The jobs and criteria were written and timestamped (2026-10-03 13:40 UTC) before the search was run on them, and they are frozen. They are never edited to raise a score. A job or criterion found to be wrong is fixed only in a new version of the set; the old version stays published, and scores are compared only within one version.

Each monthly result file carries the SHA-256 of the set it used, so a changed set shows.

Where it falls short

The ten jobs with the lowest nDCG@5 in the latest run. Full results, with the top 5 for every job: results JSON.

All jobs

Each job with its allowed level and the latest P@5. read read-only tools only, change may change data, any any level.

Show the jobs
  • Loading.

Other tool searches

The same jobs and judge can be run against any tool search that allows it. On 2026-10-03 we checked Glama: its tool search API needs an API key tied to an account, its robots.txt disallows /api/, and its terms (section 15.6) do not allow using its data to benchmark a competing directory. So no comparison is published. If another search offers open access, it will be run with the same set and shown here.

7Maps by 7IT. Also: Find a tool · Methodology · State of MCP uptime · Compare.