Benchmarks / SWE-Bench Pro
SWE-Bench Pro
Built by Scale AI · Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, et al. · released 21 Sep 2025
Long-horizon software engineering tasks taken from the commit history of 41 repositories: 11 public copyleft (GPL) repositories, 12 held-out copyleft repositories, and 18 private startup codebases. Human engineers built each environment and rewrote each task into a problem statement, a requirements brief, and an optional interface.
Frontier
61.5%
Resolve rate (public set)
Muse Spark 1.1 · harness: mini-swe-agent
9 Jul 2026 · Source: Scale AI (benchmark maintainers)
Public set (731 tasks, V1). Leaderboard shows 61.50 ± 3.10. The same model leads the private commercial set at 51.5 ± 5.5. Entry added 2026-07-09. The leaderboard lists the company as meta; open-weights status not confirmed.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Whether an agent's patch makes the new fail-to-pass tests pass without breaking the existing pass-to-pass tests. Tasks are bug fixes, features, optimizations, security updates, and UI changes in Python, Go, JavaScript, and TypeScript. Reference patches average 107.4 changed lines across 4.1 files. Scale reports separate leaderboards for the public set (731 tasks) and the private commercial set (276 tasks).
Resolve rate (public set): Share of public-set tasks where the patch passes all fail-to-pass and pass-to-pass tests. The leaderboard shows a 95% confidence interval and ranks by its upper bound. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- repo, cli
- Grading
- automated-tests
- Tasks
- 1,865
- Human reference
- Scale says tasks may take a professional software engineer hours to days. No measured human baseline is published.
- Contamination
- Public tasks come from strong-copyleft repositories, which Scale says are less likely to be in training data. The private set is not public. OpenAI's contamination probes found rarer and weaker leakage than on SWE-bench Verified, and no verbatim gold patch.
- Reuse
- Public-set tasks come from GPL-licensed repositories and stay under those licenses. The harness is MIT. The held-out and commercial sets are not released. (cite-only)
Limits to keep in mind
- OpenAI audited the public set in July 2026 and estimated that about 30% of tasks are broken (overly strict tests, underspecified prompts, low-coverage tests, misleading prompts). It withdrew its recommendation to use the benchmark. Source
- Scale released SWE-Bench Pro V2 on 2026-09-22 with 642 validated tasks and a locked offline protocol. The leaderboard still shows V1 (731-task) results, so numbers will not transfer to V2. Source
- Evaluation settings changed after launch. Launch results used a 50-turn, $2 cap. Current leaderboard runs use no cost cap and a 250-turn limit, so launch and current numbers are not comparable. Source
- Most current leaderboard entries were run with the mini-swe-agent harness, so scores reflect a fixed scaffold, not the best agent for each model. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Resolve rate (public set) | Date | Source |
|---|---|---|---|
| Muse Spark 1.1 · frontier harness: mini-swe-agent | 61.5% | 9 Jul 2026 | Scale AI · primary |
| claude-opus-4-6 (thinking) harness: mini-swe-agent | 51.9% | 8 Apr 2026 | Scale AI · primary |
| gpt-5.4 (xHigh) harness: mini-swe-agent | 59.1% | 8 Apr 2026 | Scale AI · primary |
| Claude Sonnet 4.5 harness: SWE-Agent | 43.6% | 14 Nov 2025 | Scale AI · primary |
| OpenAI GPT-5 (medium) harness: SWE-Agent | 23.3% | 21 Sep 2025 | Scale AI · primary |
Where sources disagree
Scale's public leaderboard shows a best resolve rate of 61.5% (Muse Spark 1.1, mini-swe-agent, added 2026-07-09). OpenAI's 2026-07-08 audit post says frontier models reached 80.3% on the same 731-task public split, without naming the model or harness. The gap is likely a different agent scaffold and OpenAI's own internal runs. We show the leaderboard value.
- 61.5 — labs.scale.com (primary) · Scale AI public leaderboard, Muse Spark 1.1 with mini-swe-agent, 61.50 ± 3.10, seen 2026-09-23.
- 80.3 — openai.com (lab-reported) · OpenAI post dated 2026-07-08. "Frontier models improved from a pass rate of 23.3% to 80.3% in eight months." No model or harness named.
We show 61.5. Status: open.
Timeline
Where it sits in the atlas
Go to the source
- Website scale.com
- Paper arxiv.org
- Full leaderboard labs.scale.com
- Code github.com
- Dataset huggingface.co
Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.
How to cite
Credit the original work first: SWE-Bench Pro by Scale AI (https://arxiv.org/abs/2509.16941).
Then, if you used this page:
Can Agents Work. "SWE-Bench Pro: frontier results and sources." https://canagentswork.com/benchmarks/swe-bench-pro/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-swe-bench-pro,
title = {{SWE-Bench Pro: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/swe-bench-pro/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}