Benchmarks / SWE-rebench

SWE-rebench

Built by Nebius · Ibragim Badertdinov, Alexander Golubev, Maksim Nekrashevich, Anton Shevtsov, et al. · released 26 May 2025

A continuously refreshed leaderboard of real GitHub issues. Nebius collects new issue and pull request pairs from active Python repositories every month, runs every model in the same minimal ReAct agent, and marks results where the issues predate a model's release as possibly contaminated.

Frontier

64.5%

Resolved rate

Anthropic Fable 5 [high] · harness: SWE-rebench ReAct scaffold

1 Jul 2026 · Source: Nebius (benchmark maintainers)

Rank 1 in the default window (2026-05-15 to 2026-07-01, 111 problems from 65 repositories). SEM 1.41, pass@5 78.4, $4.40 per problem. The date is the end of the time window; the leaderboard does not give a run date for each row.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Nebius

SWE-rebench: Resolved rate over time, 4 recorded results. 0%20%40%60%80%100%Jul 2025Oct 2025Jan 2026Apr 2026Jul 2026 gpt-4.1-2025-04-14: 31.1% (26 May 2025) Anthropic Opus 5 [high]: 63.4% (1 Jul 2026) Grok 4.5 [high]: 63.8% (1 Jul 2026) Anthropic Fable 5 [high]: 64.5% (1 Jul 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/swe-rebench/"><img src="https://canagentswork.com/og/benchmarks-swe-rebench.png" width="600" height="315" alt="SWE-rebench: the best result is 64.5% (Anthropic Fable 5 [high], 1 Jul 2026)." loading="lazy"></a>

Markdown:

[![SWE-rebench: the best result is 64.5% (Anthropic Fable 5 [high\], 1 Jul 2026).](https://canagentswork.com/og/benchmarks-swe-rebench.png)](https://canagentswork.com/benchmarks/swe-rebench/)

What it measures

Share of fresh issues that a model resolves, averaged over five runs, with the standard error of the mean and pass@5. Every model gets the same scaffold, the same prompts, default sampling settings, and a 128K-token context. The leaderboard shows one time window at a time; the default window on 2026-09-23 covered 111 problems from 65 repositories created between 2026-05-15 and 2026-07-01. Nebius also runs reference agents (Claude Code, Codex, Junie, Cursor) as separate rows.

Resolved rate: Mean share of problems in the selected time window resolved across five runs, with the standard error of the mean. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
repo, cli
Grading
automated-tests
Tasks
111
Contamination
Low by design. Tasks are collected after models are released, and the leaderboard flags any evaluation where issues predate the model. Older windows become contaminated over time.
Reuse
Tasks come from repositories under permissive licenses (MIT, Apache-2.0, BSD, ISC, and others that Nebius checked by hand). The leaderboard dataset and Docker images are public. (cite-only)

Limits to keep in mind

  • One fixed ReAct scaffold for all models. Nebius says model-specific tuning or another scaffold could score higher, so this measures the model in a common harness, not the best agent. Source
  • Each window is small (111 problems in the current window) and the task mix changes every month, so scores across windows are not directly comparable. Source
  • Python repositories only. Source
  • Top scores in the current window overlap within their standard errors (64.5 ± 1.41 versus 63.8 ± 0.60 and 63.4 ± 1.35), so the ranking at the top is not settled. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemResolved rateDateSource
Anthropic Opus 5 [high]
harness: SWE-rebench ReAct scaffold
63.4%1 Jul 2026Nebius · primary
Grok 4.5 [high]
harness: SWE-rebench ReAct scaffold
63.8%1 Jul 2026Nebius · primary
Anthropic Fable 5 [high] · frontier
harness: SWE-rebench ReAct scaffold
64.5%1 Jul 2026Nebius · primary
gpt-4.1-2025-04-14
harness: SWE-rebench ReAct scaffold
31.1%26 May 2025Nebius · primary

Timeline

  • 26 May 2025 — Nebius launches SWE-rebench with fresh GitHub issues. Source

Where it sits in the atlas

Software engineering

Issue resolution

Last checked 23 Sep 2026 against 5 primary sources. See an error? Tell us.

How to cite

Credit the original work first: SWE-rebench by Nebius (https://arxiv.org/abs/2505.20411).

Then, if you used this page:

Can Agents Work. "SWE-rebench: frontier results and sources." https://canagentswork.com/benchmarks/swe-rebench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-swe-rebench,
  title        = {{SWE-rebench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/swe-rebench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}