Benchmarks / APEX-Agents
APEX-Agents
Built by Mercor, Box and Harvey · Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, et al. · released 21 Jan 2026
Long, multi-application work tasks written by investment banking analysts, management consultants, and corporate lawyers. Each task sits inside a simulated project "world" with files, spreadsheets, and chat threads. An agent must find the right information and produce a client-ready output.
Frontier
73.5%
Pass@1
Claude Opus 5.5 (max) · harness: Mercor Loop agent
23 Sep 2026 · Source: Mercor (benchmark maintainers)
Rank 1 on the APEX-Agents 1.1 board, seen 2026-09-23. Pass@1 73.5% plus or minus 4.9; Mean Score 81.3% plus or minus 4.0; 950 samples. Model release date shown as 2026-09-22. Graded by an LLM judge against expert rubrics.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Whether an agent completes a professional task end to end inside a realistic workspace. Experts from top firms built 31 worlds (for example a week-long consulting project for a fictional oil and gas company) and wrote 240 tasks (80 per job) with 1 to 10 pass or fail criteria each. Version 1.1 (September 2026) tightened task specifications, added a judge that gives zero credit for hedged multiple answers, and fixed tool reliability.
Pass@1: Share of tasks where the agent passes every rubric criterion on a single attempt. Mercor also reports Mean Score, the average share of criteria passed per task. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Project level: whole projects judged by an acceptance standard
- Environment
- simulated-workplace, documents, api-tools
- Grading
- rubric, llm-judge
- Tasks
- 240
- Contamination
- Mercor says the full task set stays private so that models cannot be trained on it, while an open subset is public. The exact split between public and private tasks is not stated on the leaderboard page.
- Reuse
- CC BY 4.0 on the Hugging Face dataset. Mercor says the full task set used for the leaderboard stays private. (open-cc-by)
Limits to keep in mind
- Grading uses an LLM judge (DeepSeek-V4-Flash-0731 in v1.1) against expert rubrics, not human review of each output. Mercor reports the judge's false negative rate rose from 5.3% to 8.0% in v1.1. Source
- Version 1.1 changed tasks (480 to 240), grading, and prompts, so scores before September 2026 are not comparable with the current board. Source
- Confidence intervals are about plus or minus 5 percentage points, so the top five models overlap. Source
- Only three jobs are covered, and all worlds run in a Google Workspace style environment with Mercor's own agent loop. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Pass@1 | Date | Source |
|---|---|---|---|
| GPT-5.5 (xhigh) harness: Mercor Loop agent | 55.1% | 23 Sep 2026 | Mercor · primary |
| Claude Opus 5.5 (max) · frontier harness: Mercor Loop agent | 73.5% | 23 Sep 2026 | Mercor · primary |
| Claude Fable 5.1 (max) harness: Mercor Loop agent | 68.6% | 8 Sep 2026 | Mercor · primary |
| Gemini 3 Flash (Thinking=High) | 24% | Jan 2026 | Mercor · primary |
Timeline
Where it sits in the atlas
Finance and accountingManagement and business operations (partial)Legal (partial)
Financial analysisConsulting workLegal workProfessional deliverables
Go to the source
- Website www.mercor.com
- Paper arxiv.org
- Full leaderboard www.mercor.com
- Code github.com
- Announcement www.mercor.com
- Dataset huggingface.co
Last checked 23 Sep 2026 against 4 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: APEX-Agents by Mercor, Box, and Harvey (https://arxiv.org/abs/2601.14242).
Then, if you used this page:
Can Agents Work. "APEX-Agents: frontier results and sources." https://canagentswork.com/benchmarks/apex-agents/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-apex-agents,
title = {{APEX-Agents: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/apex-agents/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}