Benchmarks / SWE-rebench
SWE-rebench
Built by Nebius · Ibragim Badertdinov, Alexander Golubev, Maksim Nekrashevich, Anton Shevtsov, et al. · released 26 May 2025
A continuously refreshed leaderboard of real GitHub issues. Nebius collects new issue and pull request pairs from active Python repositories every month, runs every model in the same minimal ReAct agent, and marks results where the issues predate a model's release as possibly contaminated.
Frontier
64.5%
Resolved rate
Anthropic Fable 5 [high] · harness: SWE-rebench ReAct scaffold
1 Jul 2026 · Source: Nebius (benchmark maintainers)
Rank 1 in the default window (2026-05-15 to 2026-07-01, 111 problems from 65 repositories). SEM 1.41, pass@5 78.4, $4.40 per problem. The date is the end of the time window; the leaderboard does not give a run date for each row.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Share of fresh issues that a model resolves, averaged over five runs, with the standard error of the mean and pass@5. Every model gets the same scaffold, the same prompts, default sampling settings, and a 128K-token context. The leaderboard shows one time window at a time; the default window on 2026-09-23 covered 111 problems from 65 repositories created between 2026-05-15 and 2026-07-01. Nebius also runs reference agents (Claude Code, Codex, Junie, Cursor) as separate rows.
Resolved rate: Mean share of problems in the selected time window resolved across five runs, with the standard error of the mean. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- repo, cli
- Grading
- automated-tests
- Tasks
- 111
- Contamination
- Low by design. Tasks are collected after models are released, and the leaderboard flags any evaluation where issues predate the model. Older windows become contaminated over time.
- Reuse
- Tasks come from repositories under permissive licenses (MIT, Apache-2.0, BSD, ISC, and others that Nebius checked by hand). The leaderboard dataset and Docker images are public. (cite-only)
Limits to keep in mind
- One fixed ReAct scaffold for all models. Nebius says model-specific tuning or another scaffold could score higher, so this measures the model in a common harness, not the best agent. Source
- Each window is small (111 problems in the current window) and the task mix changes every month, so scores across windows are not directly comparable. Source
- Python repositories only. Source
- Top scores in the current window overlap within their standard errors (64.5 ± 1.41 versus 63.8 ± 0.60 and 63.4 ± 1.35), so the ranking at the top is not settled. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Resolved rate | Date | Source |
|---|---|---|---|
| Anthropic Opus 5 [high] harness: SWE-rebench ReAct scaffold | 63.4% | 1 Jul 2026 | Nebius · primary |
| Grok 4.5 [high] harness: SWE-rebench ReAct scaffold | 63.8% | 1 Jul 2026 | Nebius · primary |
| Anthropic Fable 5 [high] · frontier harness: SWE-rebench ReAct scaffold | 64.5% | 1 Jul 2026 | Nebius · primary |
| gpt-4.1-2025-04-14 harness: SWE-rebench ReAct scaffold | 31.1% | 26 May 2025 | Nebius · primary |
Timeline
- 26 May 2025 — Nebius launches SWE-rebench with fresh GitHub issues. Source
Where it sits in the atlas
Go to the source
- Website swe-rebench.com
- Paper arxiv.org
- Full leaderboard swe-rebench.com
- Dataset huggingface.co
Last checked 23 Sep 2026 against 5 primary sources. See an error? Tell us.
How to cite
Credit the original work first: SWE-rebench by Nebius (https://arxiv.org/abs/2505.20411).
Then, if you used this page:
Can Agents Work. "SWE-rebench: frontier results and sources." https://canagentswork.com/benchmarks/swe-rebench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-swe-rebench,
title = {{SWE-rebench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/swe-rebench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}