Benchmarks / MCPMark Verified

MCPMark Verified (MCPMark)

Built by EVAL SYS, LobeHub and NUS TRAIL · Zijian Wu, Xiangyan Liu, Xinyuan Zhang, Lingjun Chen, et al. · released 12 Jun 2026

Multi-step work on five real MCP servers: edit Notion pages and databases, run GitHub repository and CI workflows, administer and query a PostgreSQL database, browse and act on websites through Playwright, and organize files on a file system. Each task starts from a prepared state, and a script checks the end state.

Frontier

96.1%

Pass@1 (single run)

kimi-k3-max · harness: MCPMarkAgent tool-calling loop

20 Jul 2026 · Source: EVAL SYS (benchmark maintainers)

Rank 1 on the Verified board updated 2026-07-20 (122 of 127 tasks). Per service: Filesystem 93.33, GitHub 95.65, Notion 92.86, Playwright 100.00, Postgres 100.00. Single run, no interval. The board lists it as a closed model.

Saturated

The best result is at 90% or more of the ceiling.

See the full leaderboard at EVAL SYS

MCPMark Verified: Pass@1 (single run) over time, 2 recorded results. 0%20%40%60%80%100%Jun 2026Jul 2026Aug 2026 gpt-5-5-xhigh: 92.9% (12 Jun 2026) kimi-k3-max: 96.1% (20 Jul 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/mcpmark-verified/"><img src="https://canagentswork.com/og/benchmarks-mcpmark-verified.png" width="600" height="315" alt="MCPMark Verified: the best result is 96.1% (kimi-k3-max, 20 Jul 2026)." loading="lazy"></a>

Markdown:

[![MCPMark Verified: the best result is 96.1% (kimi-k3-max, 20 Jul 2026).](https://canagentswork.com/og/benchmarks-mcpmark-verified.png)](https://canagentswork.com/benchmarks/mcpmark-verified/)

What it measures

Whether an agent can carry out create, read, update, and delete work through MCP tools and leave the system in the required state. 127 tasks: Filesystem 30, Notion 28, Playwright 25, GitHub 23, Postgres 21. Examples: set up an ESLint workflow on all pull requests, write a PostgreSQL function for inventory transfers with audit logging, recolor elements on a Notion page, extract contact details from mixed file formats. The Verified set (June 2026) is a subset of the standard tasks that pins every server version (server-filesystem 2025.12.18, github-mcp-server v0.15.0, notion-mcp-server 1.9.1, playwright/mcp 0.0.68, postgres-mcp 0.3.0) and stabilizes every verifier.

Pass@1 (single run): Share of the 127 tasks whose verification script passes on one run (run-1). A completion rate. The Legacy board also reported Pass@4 and Pass^4 over four runs; the Verified board is single-run. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
api-tools, browser, repo
Grading
state-check, automated-tests
Tasks
127
Contamination
All tasks and verifiers are public on GitHub. GitHub tasks recreate real open-source repositories from state templates.
Reuse
Tasks, initial states, and verifiers ship in the GitHub repository under Apache License 2.0. (open-apache)

Limits to keep in mind

  • The top three entries exceed 90% on the Verified set, so it separates the best systems little. Source
  • Single-run scores with no confidence interval. The board notes that kimi-k2-7-code's filesystem score comes from its second recorded run, and that claude-opus-4-8-max and kimi-k2-6 are scores-only entries without per-task results. Source
  • Results before the Verified release (June 2026) used unpinned server versions and older verifiers; the maintainers call them deprecated and not comparable. Source
  • Only five MCP servers, and the tasks were written by the benchmark's own team with AI help rather than sampled from real workloads. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemPass@1 (single run)DateSource
kimi-k3-max · frontier
harness: MCPMarkAgent tool-calling loop
96.1%20 Jul 2026EVAL SYS · primary
gpt-5-5-xhigh
harness: MCPMarkAgent tool-calling loop
92.9%12 Jun 2026EVAL SYS · primary

Timeline

  • 20 Jul 2026 — MCPMark Verified saturates: kimi-k3-max passes 96.06%, three models above 90%. Source
  • 12 Jun 2026 — MCPMark Verified becomes the default set; gpt-5.5 (xhigh) passes 92.9%. Source

Where it sits in the atlas

Work ladder: Context only. It spans many occupations, so it does not set the step of a job family. How the ladder works

Office and administrative supportSoftware engineering (partial)Data and analytics (partial)

MCP tool useEnterprise workflowsData engineering and SQL

Last checked 24 Sep 2026 against 8 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: MCPMark Verified by EVAL SYS, LobeHub, and NUS TRAIL (https://arxiv.org/abs/2509.24002).

Then, if you used this page:

Can Agents Work. "MCPMark Verified: frontier results and sources." https://canagentswork.com/benchmarks/mcpmark-verified/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-mcpmark-verified,
  title        = {{MCPMark Verified: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/mcpmark-verified/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}