Benchmark Hall
42 benchmarks, each built by the institutions named on its page. We show the current frontier result and link to the full leaderboard at the source. Grouped by status: from open problems to benchmarks that the frontier has already passed.
By function: Issue resolution 6Feature development 5Long-horizon autonomy 5Incident response 4Professional deliverables 4Shipping pull requests 3Terminal operations 3Vulnerability research 3Web research 3Codebase evolution 2Computer use 2Customer service 2Enterprise workflows 2Financial analysis 2Freelance projects 2Research replication 2Spreadsheet work 2Asking for clarification 1Autonomous R&D 1Clinical records work 1Code review 1Compliance operations 1Consulting work 1CRM operations 1Data engineering and SQL 1FinOps 1Infrastructure as code 1Legal work 1ML engineering 1Performance engineering 1Running a business 1Test integrity 1
Open 2
The best result is below 20% of the ceiling, or below half of human parity.
Emerging 8
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
- 36.7%GPT-4 with retrieval-augmented generation · 2024Emerging
- 47%Claude Opus 4.7 (Adaptive Reasoning, Max Effort) · 27 May 2026Emerging
- 42.9%TTE-MatrixAgent + DeepSeek-V3.2 · 10 Nov 2025Emerging
Strong 13
The best result is at 60–90% of the ceiling, or at or above human parity.
Unrated 7
No fixed reference point (for example Elo scores or field signals).
- 1385 EloClaude Sonnet 4.5 · 3 Nov 2025Unrated
- 1846 EloClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) · 23 Sep 2026Unrated
No results yet 1
We have not recorded a result yet.
Saturated 8
The best result is at 90% or more of the ceiling.
- OSWorld-Verifiedby HKU, Salesforce AI Research, Carnegie Mellon University and University of Waterloo90.2%Intelligence-Indeed Agent · 25 Jul 2026Saturated
- 96.7%Genloop's Sentinel Agent v2 Pro · 1 Mar 2026Saturated
Solved 1
Its maintainers or a major evaluator declared it solved.
- 95.5%Claude Code + Claude Opus 4.5 (after manual regrading) · 3 Dec 2025Solved
Retired 2
Maintainers or a major user stopped using it as a frontier measure.
- 79.2%Sonar Foundation Agent + Claude 4.5 Opus · 5 Dec 2025Retired
- 84.7%NexAU-AHE + GPT-5.5 · 23 Apr 2026Retired
Status rules are on the method page.