Benchmarks / MLE-bench

MLE-bench

Built by OpenAI · Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, et al. · released 9 Oct 2024

75 Kaggle competitions rebuilt offline. An agent gets the data and description, trains models, and submits predictions. Its score is compared with the real human leaderboard to see whether it would have won a bronze, silver, or gold medal.

Frontier

64.4%

Any medal (%)

Famou-Agent 2.0 (Gemini-3-Pro-Preview) · agent: Famou-Agent 2.0 (Baidu)

23 Feb 2026 · Source: OpenAI (benchmark maintainers)

64.44 ± 1.18 over 24 hours per competition. Top of the main leaderboard table. Disarray reports 77.78 but used test-set feedback and is listed separately as not comparable.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at OpenAI

MLE-bench: Any medal (%) over time, 5 recorded results. 0%20%40%60%80%100%Oct 2024Jan 2025Apr 2025Jul 2025Oct 2025Jan 2026Apr 2026 AIDE + o1-preview: 17.1% (8 Oct 2024) R&D-Agent (o3 + GPT-4.1): 30.2% (15 Aug 2025) Famou-Agent 2.0 (Gemini-2.5-Pro): 59.6% (27 Dec 2025) Famou-Agent 2.0 (Gemini-3-Pro-Preview): 64.4% (23 Feb 2026) AIBuildAI (Claude Opus 4.6): 63.1% (6 Mar 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/mle-bench/"><img src="https://canagentswork.com/og/benchmarks-mle-bench.png" width="600" height="315" alt="MLE-bench: the best result is 64.4% (Famou-Agent 2.0 (Gemini-3-Pro-Preview), 23 Feb 2026)." loading="lazy"></a>

Markdown:

[![MLE-bench: the best result is 64.4% (Famou-Agent 2.0 (Gemini-3-Pro-Preview), 23 Feb 2026).](https://canagentswork.com/og/benchmarks-mle-bench.png)](https://canagentswork.com/benchmarks/mle-bench/)

What it measures

Whether an agent can do end-to-end machine learning engineering: prepare data, train and tune models, and produce a valid submission. The headline is the share of competitions where the agent's best submission reaches at least a bronze medal, averaged over at least 3 seeds on all 75 competitions. The repo also reports Low (22 competitions, the "Lite" set), Medium, and High splits.

Any medal (%): Share of the 75 competitions where the agent earns any medal, reported as the mean over seeds. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Project level: whole projects judged by an acceptance standard
Environment
cli
Grading
outcome-metric
Tasks
75
Human reference
Medal thresholds come from the public Kaggle leaderboards, so a medal means the agent beat most human participants in that competition.
Contamination
The competitions are public and top solutions are online. The paper ran familiarity experiments and found no clear link between a model's familiarity with a competition and its score.
Reuse
Code is MIT licensed. Competition data comes from Kaggle and users must accept each competition's rules to download it. (cite-only)

Limits to keep in mind

  • Since 2026-04-24 the maintainers accept no new leaderboard submissions while they design a fairer submission process. Source
  • Two submissions (Disarray 77.78%, LoongFlow 62.66%) are listed separately because they used test-set feedback and are not comparable with the main table. Source
  • Known grading issues in several competitions are left unfixed to keep the v1 leaderboard comparable. Fixes are planned for a v2 release. Source
  • Agents run for 24 hours per competition with modern models, while Kaggle participants worked under different conditions, so a medal is not a like-for-like human comparison. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemAny medal (%)DateSource
AIBuildAI (Claude Opus 4.6)63.1%6 Mar 2026OpenAI · primary
Famou-Agent 2.0 (Gemini-3-Pro-Preview) · frontier64.4%23 Feb 2026OpenAI · primary
Famou-Agent 2.0 (Gemini-2.5-Pro)59.6%27 Dec 2025OpenAI · primary
R&D-Agent (o3 + GPT-4.1)
harness: R&D-Agent
30.2%15 Aug 2025OpenAI · primary
AIDE + o1-preview
harness: AIDE
17.1%8 Oct 2024OpenAI · primary

Timeline

  • 24 Apr 2026 — MLE-bench pauses new leaderboard submissions. Source
  • 23 Feb 2026 — MLE-bench leader earns medals in 64.44% of competitions. Source
  • 9 Oct 2024 — OpenAI releases MLE-bench with 75 Kaggle competitions. Source

Where it sits in the atlas

ML and research engineeringData and analytics (partial)

ML engineering

Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.

How to cite

Credit the original work first: MLE-bench by OpenAI (https://arxiv.org/abs/2410.07095).

Then, if you used this page:

Can Agents Work. "MLE-bench: frontier results and sources." https://canagentswork.com/benchmarks/mle-bench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-mle-bench,
  title        = {{MLE-bench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/mle-bench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}