Benchmarks / RE-Bench
RE-Bench
Built by METR · Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, et al. · released 22 Nov 2024
Seven open-ended ML research engineering environments built from scratch by METR, for example writing a faster GPU kernel, fixing a corrupted model embedding, or inferring a scaling law. Humans and agents get the same machine, GPUs, and scoring function, and try to push the score as high as they can in a fixed time.
Frontier
37%
Human-expert percentile matched (8-hour budget)
Claude 3.5 Sonnet (New) (Modular scaffold) · harness: Modular
31 Jan 2025 · Source: METR (benchmark maintainers)
METR says Claude 3.5 Sonnet performed comparable to a 37th-percentile human expert at 8 hours per task. This is the latest 8-hour, 7-task figure METR has published. Newer models were evaluated on a 5-task subset at a 32-hour budget and are not comparable.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
How an agent's research engineering output compares with human experts under the same conditions. Each environment's raw score is normalized so the starting solution is 0 and METR's reference solution is 1. METR reports the average normalized score by total time budget, using best-of-k over shorter runs for agents. The comparison set is 71 eight-hour attempts by 61 ML experts, whose average normalized score was 0.64. We track the human-expert percentile that METR assigns to an agent at an 8-hour budget.
Human-expert percentile matched (8-hour budget): Percentile of METR's human expert 8-hour attempts that the agent's average normalized score matches when the agent also gets an 8-hour total budget. 50 means the median expert. Higher is better.
Status compares the frontier with human parity at 50%. Median human expert attempt over 8 hours.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli
- Grading
- outcome-metric
- Tasks
- 7
- Human reference
- 71 eight-hour attempts by 61 human experts (METR hiring applicants, professional network, and graduate students). Average normalized score 0.64; 82% of attempts scored above zero and 24% matched or beat the reference solution.
- Contamination
- Environments were created from scratch and had not appeared in training data at launch. Public discussion of solutions since then may contaminate future runs. METR keeps other environments held out.
- Reuse
- Repo is MIT licensed. Reference solutions are password-protected. METR asks users to keep the tasks out of training data and not to publish solutions. (open-mit)
Limits to keep in mind
- Only 7 environments. METR's later reports often use a 5-task subset and fold RE-Bench into time-horizon estimates, so there is no single up-to-date leaderboard number. Source
- Agents beat humans at a 2-hour budget (about 4 times the human score) but humans pull ahead at 8 hours and reach about twice the best agent at 32 hours, so the headline depends on the budget. Source
- Agents can run the scoring function at will, and METR found reward hacking, for example faking a training run's output. Detected hacks are scored as failures. Source
- Environments have clear goals and fast feedback, unlike much real research, and human scores vary a lot by recruiting source (0.48 for hiring applicants, 0.98 for professional network). Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
- 22 Nov 2024 — METR releases RE-Bench with 71 human expert baselines. Source
Where it sits in the atlas
Go to the source
- Website metr.org
- Paper arxiv.org
- Code github.com
- Announcement metr.org
Last checked 23 Sep 2026 against 9 primary sources. See an error? Tell us.
How to cite
Credit the original work first: RE-Bench by METR (https://arxiv.org/abs/2411.15114).
Then, if you used this page:
Can Agents Work. "RE-Bench: frontier results and sources." https://canagentswork.com/benchmarks/re-bench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-re-bench,
title = {{RE-Bench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/re-bench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}