Benchmarks / RE-Bench

RE-Bench

Built by METR · Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, et al. · released 22 Nov 2024

Seven open-ended ML research engineering environments built from scratch by METR, for example writing a faster GPU kernel, fixing a corrupted model embedding, or inferring a scaling law. Humans and agents get the same machine, GPUs, and scoring function, and try to push the score as high as they can in a fixed time.

Frontier

37%

Human-expert percentile matched (8-hour budget)

Claude 3.5 Sonnet (New) (Modular scaffold) · harness: Modular

31 Jan 2025 · Source: METR (benchmark maintainers)

METR says Claude 3.5 Sonnet performed comparable to a 37th-percentile human expert at 8 hours per task. This is the latest 8-hour, 7-task figure METR has published. Newer models were evaluated on a 5-task subset at a 32-hour budget and are not comparable.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

RE-Bench: Human-expert percentile matched (8-hour budget) over time, 2 recorded results. 0%20%40%60%Feb 2025human parity o1 (AIDE scaffold): 30% (31 Jan 2025) Claude 3.5 Sonnet (New) (Modular scaffold): 37% (31 Jan 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/re-bench/"><img src="https://canagentswork.com/og/benchmarks-re-bench.png" width="600" height="315" alt="RE-Bench: the best result is 37% (Claude 3.5 Sonnet (New) (Modular scaffold), 31 Jan 2025)." loading="lazy"></a>

Markdown:

[![RE-Bench: the best result is 37% (Claude 3.5 Sonnet (New) (Modular scaffold), 31 Jan 2025).](https://canagentswork.com/og/benchmarks-re-bench.png)](https://canagentswork.com/benchmarks/re-bench/)

What it measures

How an agent's research engineering output compares with human experts under the same conditions. Each environment's raw score is normalized so the starting solution is 0 and METR's reference solution is 1. METR reports the average normalized score by total time budget, using best-of-k over shorter runs for agents. The comparison set is 71 eight-hour attempts by 61 ML experts, whose average normalized score was 0.64. We track the human-expert percentile that METR assigns to an agent at an 8-hour budget.

Human-expert percentile matched (8-hour budget): Percentile of METR's human expert 8-hour attempts that the agent's average normalized score matches when the agent also gets an 8-hour total budget. 50 means the median expert. Higher is better.

Status compares the frontier with human parity at 50%. Median human expert attempt over 8 hours.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli
Grading
outcome-metric
Tasks
7
Human reference
71 eight-hour attempts by 61 human experts (METR hiring applicants, professional network, and graduate students). Average normalized score 0.64; 82% of attempts scored above zero and 24% matched or beat the reference solution.
Contamination
Environments were created from scratch and had not appeared in training data at launch. Public discussion of solutions since then may contaminate future runs. METR keeps other environments held out.
Reuse
Repo is MIT licensed. Reference solutions are password-protected. METR asks users to keep the tasks out of training data and not to publish solutions. (open-mit)

Limits to keep in mind

  • Only 7 environments. METR's later reports often use a 5-task subset and fold RE-Bench into time-horizon estimates, so there is no single up-to-date leaderboard number. Source
  • Agents beat humans at a 2-hour budget (about 4 times the human score) but humans pull ahead at 8 hours and reach about twice the best agent at 32 hours, so the headline depends on the budget. Source
  • Agents can run the scoring function at will, and METR found reward hacking, for example faking a training run's output. Detected hacks are scored as failures. Source
  • Environments have clear goals and fast feedback, unlike much real research, and human scores vary a lot by recruiting source (0.48 for hiring applicants, 0.98 for professional network). Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemHuman-expert percentile matched (8-hour budget)DateSource
o1 (AIDE scaffold)
harness: AIDE
30%31 Jan 2025METR · primary
Claude 3.5 Sonnet (New) (Modular scaffold) · frontier
harness: Modular
37%31 Jan 2025METR · primary

Timeline

  • 22 Nov 2024 — METR releases RE-Bench with 71 human expert baselines. Source

Where it sits in the atlas

ML and research engineering

Autonomous R&DLong-horizon autonomyPerformance engineering

Last checked 23 Sep 2026 against 9 primary sources. See an error? Tell us.

How to cite

Credit the original work first: RE-Bench by METR (https://arxiv.org/abs/2411.15114).

Then, if you used this page:

Can Agents Work. "RE-Bench: frontier results and sources." https://canagentswork.com/benchmarks/re-bench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-re-bench,
  title        = {{RE-Bench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/re-bench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}