Benchmarks / METR Task-Completion Time Horizons
METR Task-Completion Time Horizons (METR Time Horizon)
Built by METR · Thomas Kwa, Ben West, Joel Becker, Amy Deng, et al. · released 18 Mar 2025
METR runs AI agents on software, ML, and cybersecurity tasks that skilled humans were timed on, then fits a curve of success against human task length. The 50% time horizon is the human task length at which the agent succeeds half the time. METR reports the frontier horizon has doubled about every 7 months since 2019, and about every 4 months since 2023 under the TH1.1 suite.
Frontier
17 h
50% task-completion time horizon
Claude Mythos Preview (early) · harness: METR ReAct agent (Inspect)
7 Apr 2026 · Source: METR (benchmark maintainers)
About 17.4 hours. Added to the page 2026-05-08 together with a notice that measurements above 16 hours are unreliable with the current task suite. 80% horizon 185.9 minutes. Date is the release_date field in METR's data file.
No fixed reference point (for example Elo scores or field signals).
What it measures
How long a task (in expert human time) an agent can complete with 50% reliability. The Time Horizon 1.1 suite has 228 tasks from HCAST, RE-Bench, and short SWAA tasks, with 31 tasks estimated at 8 hours or more for humans. Human times come from contracted professionals with about 5 years of experience, given the same instructions and tools as the agents. Each model runs with a scaffold that METR chooses after a small elicitation phase; runs are checked for reward hacking. METR also reports an 80% horizon, which is several times shorter.
50% task-completion time horizon: Human task length at which the fitted logistic curve predicts a 50% success rate for the agent. Time Horizon 1.1 suite unless noted. Higher is better.
No fixed reference point, so we do not rate its status. Open-ended scale in human minutes. METR says measurements above 16 hours (960 minutes) are unreliable with the current suite.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli, repo
- Grading
- automated-tests, state-check
- Tasks
- 228
- Human reference
- Skilled contractors (software, ML, cybersecurity; about 5 years of experience) attempt each task; the geometric mean of successful completion times sets the task length. Only 5 of the 31 tasks of 8 or more hours have measured human baselines; the rest use estimates.
- Contamination
- Most tasks are private. SWAA short tasks and some HCAST tasks are described publicly. METR screens runs for reward hacking by keyword search, LLM flags, and human review.
- Reuse
- Horizon estimates and run data are public in eval-analysis-public (no license file seen). Many tasks are private. (cite-only)
Limits to keep in mind
- Measurements above 16 hours are unreliable with the current task suite (notice added 2026-05-08). The top model, Claude Mythos Preview (early), is at 1,044.8 minutes, above that line, with a 95% CI of 508.9 to 3,304.3 minutes. Source
- Tasks are self-contained and well specified, so the horizon is closer to what a low-context new hire or contractor could do, not a professional with full context. METR says agents do worse on messier tasks and when graded holistically. Source
- Task composition changes the trend. Moving from TH1 to TH1.1 changed the post-2023 doubling time from 165 to 131 days and moved individual estimates by up to 57%. A regularization fix on 2026-03-03 changed estimates again (Claude Opus 4.5 went from 320 to 293 minutes). Source
- Coverage is not complete. As of 2026-09-23 METR listed Claude Opus 4.7, Grok 4.3, and GPT-5.5 as recent models without time horizons. Evaluations take 1 to 2 weeks or more. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | 50% task-completion time horizon | Date | Source |
|---|---|---|---|
| Claude Mythos Preview (early) · frontier harness: METR ReAct agent (Inspect) | 17 h | 7 Apr 2026 | METR · primary |
| Claude Opus 4.6 harness: METR ReAct agent (Inspect) | 12 h | 5 Feb 2026 | METR · primary |
| Claude Opus 4.5 harness: METR ReAct agent (Inspect) | 4.9 h | 24 Nov 2025 | METR · primary |
| GPT-5 harness: Triframe (Inspect) | 3.4 h | 7 Aug 2025 | METR · primary |
| Claude 3.7 Sonnet harness: METR ReAct agent (Inspect) | 1 h | 24 Feb 2025 | METR · primary |
| GPT-4 (0314) harness: modular-public | 4 min | 14 Mar 2023 | METR · primary |
Where sources disagree
METR's Time Horizon 1.1 blog post (2026-01-29) gives Claude Opus 4.5 a 50% horizon of 320 minutes. METR's live data file gives 293.0 minutes. Both are METR sources. The results page lists a 2026-03-03 update that "corrected a regularization mistake that affected our measurements", which explains the change. GPT-5 moved the same way (214 to 203 minutes).
- 293 — metr.org (primary) · Live data file linked from metr.org/time-horizons, retrieved 2026-09-23. CI 161.7 to 623.7.
- 320 — metr.org (primary) · TH1.1 launch post, table "Changes to Model Horizon Estimates". CI 170 to 729.
We show 293. Status: resolved.
Timeline
Where it sits in the atlas
Software engineeringML and research engineering (partial)Security (partial)
Go to the source
- Website metr.org
- Paper arxiv.org
- Full leaderboard metr.org
- Code github.com
- Announcement metr.org
- Dataset metr.org
Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.
How to cite
Credit the original work first: METR Task-Completion Time Horizons by METR (https://arxiv.org/abs/2503.14499).
Then, if you used this page:
Can Agents Work. "METR Task-Completion Time Horizons: frontier results and sources." https://canagentswork.com/benchmarks/metr-time-horizons/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-metr-time-horizons,
title = {{METR Task-Completion Time Horizons: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/metr-time-horizons/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}