Jobs / ML and research engineering
Can agents do ML and research engineering work?
Train and evaluate models, run experiments, and reproduce research.
- US jobs
- 37,200 (0.0% of employment)
- Typical wage
- $140K a year (employment-weighted mean of occupation medians)
- Occupations
- 1 in O*NET 31.0
- Evidence
- 6 benchmarks from 5 institutions
Task level
Partial
Mixed, and partly stale. On RE-Bench, METR's AI research engineering tasks, the best recorded agent matched the 37th percentile of human experts at an 8-hour budget, but that result is from January 2025. CORE-Bench, which asks agents to reproduce published results from the authors' code, was declared solved in December 2025 (95.5% after manual regrading).
Evidence: RE-Bench, CORE-Bench, METR Task-Completion Time Horizons, Terminal-Bench 2.0 — by Laude Institute, METR, Princeton University, and Stanford University. Reading dated 23 Sep 2026.
Agents succeed on some scoped tasks. Reliable or end-to-end work is not shown.
Project level
Partial · thin evidence (fewer than 2 independent institutions)
Mixed. On MLE-bench, 75 offline Kaggle competitions, the best agent wins a medal in 64.4% of them. On PaperBench, where agents replicate ML papers from scratch, the best official result is 26.0% (o1, April 2025); on a small subset, ML PhDs scored 41.4%. PaperBench has no newer official results.
Evidence: MLE-bench, PaperBench — by OpenAI. Reading dated 23 Sep 2026.
Agents succeed on some scoped tasks. Reliable or end-to-end work is not shown.
Benchmarks for this work
Also relevant
These benchmarks cover part of this family's work.
- 95.5%Claude Code + Claude Opus 4.5 (after manual regrading) · 3 Dec 2025Solved
- 84.7%NexAU-AHE + GPT-5.5 · 23 Apr 2026Retired
Functions: Research replicationLong-horizon autonomyML engineeringAutonomous R&DPerformance engineeringTerminal operationsFeature development
What we do not know
- Benchmarks measure tasks, not whole jobs. A job is a bundle of tasks, judgment, and relationships.
- Capability is not adoption. A task that agents can do in a benchmark may still be done by people at work.
- Most results come from single runs. Reliability over many runs is often unknown.
Occupations in this family
Sorted by US employment (BLS OEWS May 2025). O*NET sub-occupations share the employment figure of their SOC group.
| Occupation | SOC | US jobs |
|---|---|---|
| Computer and Information Research Scientists | 15-1221.00 | 37K |
How to cite
Credit the original work first: the benchmarks listed on this page, and their institutions.
Then, if you used this page:
Can Agents Work. "Can agents do ML and research engineering work?." https://canagentswork.com/jobs/ml-research-engineering/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-jobs-ml-research-engineering,
title = {{Can agents do ML and research engineering work?}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/jobs/ml-research-engineering/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}