Benchmarks / MCP Atlas
MCP Atlas (MCP-Atlas)
Built by Scale AI · released 19 Sep 2025
Single-turn requests that an agent can only answer by chaining tools on real Model Context Protocol (MCP) servers: web search, maps, weather, file systems, MongoDB and Airtable databases, Notion, Slack, email, arXiv and PubMed, market data, Git and GitHub, and code runners. Each request needs 3 to 6 tool calls, usually across several servers.
Frontier
88.1%
Pass rate
Muse Spark 1.1
9 Jul 2026 · Source: Scale AI (benchmark maintainers)
Rank 1 on the board seen 2026-09-24, plus or minus 1.95. Fable 5.1 (added 2026-09-04) shares rank 1 at 87.2% plus or minus 2.05; claude-opus-5 (xhigh) is at 85.8%. The page header still says 83.6%.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Whether a model can find the right tools in a noisy menu, call them correctly, recover from errors, and combine the outputs into a correct answer. 1,000 human-written tasks (500 public, 500 private) run against 36 real MCP servers in Docker with fixed data, grouped by Scale as search and fetch (brave_search, ddg_search, exa, fetch, weather, google-maps), analytics (mongodb, airtable, calculator), productivity (filesystem, notion, slack, google-workspace, arxiv, pubmed), financial (twelvedata, alchemy), and coding (git, github, mcp-code-executor, cli-mcp-server, e2b-server). Each task exposes 10 to 25 tools, of which 3 to 7 are needed and the rest are distractors.
Pass rate: Share of all 1,000 tasks where the final answer covers at least 75% of the ground-truth claims. An LLM judge scores each claim 0, 0.5, or 1; coverage is the mean, and a task passes at 0.75 or above. A completion rate at the task level, with partial credit inside the threshold. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- api-tools
- Grading
- llm-judge, rubric
- Tasks
- 1,000
- Contamination
- Half of the tasks are held out. Stateful servers are seeded with fixed data so answers stay stable; search tasks hit live services.
- Reuse
- The 500 public tasks are on Hugging Face under CC BY 4.0 (prompts, enabled tools, claims, and reference trajectories). The other 500 tasks are private. (open-cc-by)
Limits to keep in mind
- The page's header ("83.6% Top Pass Rate") and Key Findings ("Top performer: Claude Opus 4.5 with 62.3%") lag the leaderboard table, which shows 88.1% for Muse Spark 1.1. Source
- In April 2026 Scale changed the judge, added retries for transient tool errors, and replaced the 20-turn limit with a budget of 100 tool calls, then re-scored every model. Scores from before the change are not comparable. Source
- Grading uses an LLM judge (Gemini-2.5-Pro on the page; the open harness defaults to gemini-3.1-pro-preview) against claim lists. A task can pass with a quarter of its claims wrong. Source
- Tasks are single-turn lookups and computations with a known answer, not open-ended work products. Coverage is fixed to these 36 servers (maintainers' list): airtable, alchemy, arxiv, brave-search, calculator, cli-mcp-server, clinicaltrialsgov-mcp-server, context7, ddg-search, desktop-commander, e2b-server, exa, fetch, filesystem, git, github, google-maps, google-workspace, lara-translate, mcp-code-executor, mcp-server-code-runner, memory, met-museum, mongodb, national-parks, notion, open-library, osm-mcp-server, oxylabs, pubmed, slack, twelvedata, weather, weather-data, whois, wikipedia. The page counts 220 tools; the maintainers' list counts 307 tools on the same 36 servers. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
Where it sits in the atlas
Work ladder: Context only. It spans many occupations, so it does not set the step of a job family. How the ladder works
Office and administrative supportData and analytics (partial)Software engineering (partial)Finance and accounting (partial)
Go to the source
- Website labs.scale.com
- Paper arxiv.org
- Full leaderboard labs.scale.com
- Code github.com
- Announcement scale.com
- Dataset huggingface.co
Last checked 24 Sep 2026 against 6 primary sources. See an error? Tell us.
How to cite
Credit the original work first: MCP Atlas by Scale AI (https://arxiv.org/abs/2602.00933).
Then, if you used this page:
Can Agents Work. "MCP Atlas: frontier results and sources." https://canagentswork.com/benchmarks/mcp-atlas/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-mcp-atlas,
title = {{MCP Atlas: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/mcp-atlas/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}