Benchmarks / CORE-Bench

CORE-Bench

Built by Princeton University · Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl, et al. · released 17 Sep 2024

Computational reproducibility tasks from 90 published papers in computer science, social science, and medicine, all sourced from CodeOcean capsules known to run. An agent gets the paper's code and data, must install dependencies, run the code, and answer questions about the results, including values read from figures.

Frontier

95.5%

CORE-Bench-Hard accuracy (public test set)

Claude Code + Claude Opus 4.5 (after manual regrading) · harness: Claude Code

3 Dec 2025 · Source: Princeton University (benchmark maintainers)

HAL regraded 8 tasks by hand and removed 1 task with a dead data URL. The agent failed 2 remaining tasks. HAL declared the benchmark solved on this basis.

Solved

Its maintainers or a major evaluator declared it solved.

Solved on 3 Dec 2025: HAL (Princeton) declared CORE-Bench solved after Claude Opus 4.5 in a Claude Code scaffold scored 77.78% on Hard, rising to 95.5% after HAL fixed grading errors in 8 tasks by manual scoring and removed 1 task whose data URL had died. HAL said it would open a private test set of new papers. Source.

See the full leaderboard at Princeton University

CORE-Bench: CORE-Bench-Hard accuracy (public test set) over time, 5 recorded results. 0%20%40%60%80%100%Oct 2024Jan 2025Apr 2025Jul 2025Oct 2025 CORE-Agent + GPT-4o: 21.5% (17 Sep 2024) CORE-Agent + Claude Opus 4.1: 51.1% (2025) Claude Code + Claude Sonnet 4.5: 62.2% (2025) Claude Code + Claude Opus 4.5 (automated grading): 77.8% (3 Dec 2025) Claude Code + Claude Opus 4.5 (after manual regrading): 95.5% (3 Dec 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/core-bench/"><img src="https://canagentswork.com/og/benchmarks-core-bench.png" width="600" height="315" alt="CORE-Bench: the best result is 95.5% (Claude Code + Claude Opus 4.5 (after manual regrading), 3 Dec 2025)." loading="lazy"></a>

Markdown:

[![CORE-Bench: the best result is 95.5% (Claude Code + Claude Opus 4.5 (after manual regrading), 3 Dec 2025).](https://canagentswork.com/og/benchmarks-core-bench.png)](https://canagentswork.com/benchmarks/core-bench/)

What it measures

Whether an agent can reproduce a paper's reported results from its own repository. Three levels: Easy gives the outputs, Medium gives a Docker command, Hard gives only the codebase. The headline is accuracy on CORE-Bench-Hard over the 45-paper public test set, where a task counts only if every question is answered within the tolerance set from three manual runs.

CORE-Bench-Hard accuracy (public test set): Share of the 45 Hard test-set tasks where the agent answers every task question correctly. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli, repo
Grading
automated-tests
Tasks
45
Human reference
Every task was reproduced by hand three times to set answer tolerances, so a correct answer means matching what a careful human reproduction produced.
Contamination
The 45 test papers are public and named. HAL keeps a second private set of papers, disclosed only after the 80% threshold was crossed, for future evaluation.
Reuse
Harness and code are MIT licensed. Papers and capsules come from CodeOcean and keep their own licenses. (open-mit)

Limits to keep in mind

  • HAL found grading errors in 9 tasks that only surfaced with strong agents. Deterministic outputs were penalized for tiny floating-point differences, and some tasks were underspecified. Source
  • Papers were filtered to those that run in under 45 minutes, and the agent reproduces only selected results, so real reproduction is harder than the benchmark. Source
  • Scores depend heavily on scaffold. Opus 4.5 scored 42.22% with HAL's CORE-Agent and 77.78% with Claude Code before any regrading. Source
  • HAL has paused adding new models to its leaderboards while it focuses on reliability work. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemCORE-Bench-Hard accuracy (public test set)DateSource
Claude Code + Claude Opus 4.5 (automated grading)
harness: Claude Code
77.8%3 Dec 2025Princeton University · primary
Claude Code + Claude Opus 4.5 (after manual regrading) · frontier
harness: Claude Code
95.5%3 Dec 2025Princeton University · primary
CORE-Agent + Claude Opus 4.1
harness: CORE-Agent
51.1%2025Princeton University · primary
Claude Code + Claude Sonnet 4.5
harness: Claude Code
62.2%2025Princeton University · primary
CORE-Agent + GPT-4o
harness: CORE-Agent
21.5%17 Sep 2024Princeton University · primary

Timeline

  • 3 Dec 2025 — HAL declares CORE-Bench solved. Source
  • 17 Sep 2024 — Princeton releases CORE-Bench for computational reproducibility. Source

Where it sits in the atlas

Science and researchML and research engineering (partial)

Research replication

Last checked 23 Sep 2026 against 5 primary sources. See an error? Tell us.

How to cite

Credit the original work first: CORE-Bench by Princeton University (https://arxiv.org/abs/2409.11363).

Then, if you used this page:

Can Agents Work. "CORE-Bench: frontier results and sources." https://canagentswork.com/benchmarks/core-bench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-core-bench,
  title        = {{CORE-Bench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/core-bench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}