Benchmark Hall

42 benchmarks, each built by the institutions named on its page. We show the current frontier result and link to the full leaderboard at the source. Grouped by status: from open problems to benchmarks that the frontier has already passed.

By function: Issue resolution 6Feature development 5Long-horizon autonomy 5Incident response 4Professional deliverables 4Shipping pull requests 3Terminal operations 3Vulnerability research 3Web research 3Codebase evolution 2Computer use 2Customer service 2Enterprise workflows 2Financial analysis 2Freelance projects 2Research replication 2Spreadsheet work 2Asking for clarification 1Autonomous R&D 1Clinical records work 1Code review 1Compliance operations 1Consulting work 1CRM operations 1Data engineering and SQL 1FinOps 1Infrastructure as code 1Legal work 1ML engineering 1Performance engineering 1Running a business 1Test integrity 1

Open 2

The best result is below 20% of the ceiling, or below half of human parity.

Emerging 8

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

Strong 13

The best result is at 60–90% of the ceiling, or at or above human parity.

Unrated 7

No fixed reference point (for example Elo scores or field signals).

No results yet 1

We have not recorded a result yet.

Saturated 8

The best result is at 90% or more of the ceiling.

Solved 1

Its maintainers or a major evaluator declared it solved.

Retired 2

Maintainers or a major user stopped using it as a frontier measure.

Status rules are on the method page.