Benchmarks / METR Task-Completion Time Horizons

METR Task-Completion Time Horizons (METR Time Horizon)

Built by METR · Thomas Kwa, Ben West, Joel Becker, Amy Deng, et al. · released 18 Mar 2025

METR runs AI agents on software, ML, and cybersecurity tasks that skilled humans were timed on, then fits a curve of success against human task length. The 50% time horizon is the human task length at which the agent succeeds half the time. METR reports the frontier horizon has doubled about every 7 months since 2019, and about every 4 months since 2023 under the TH1.1 suite.

Frontier

17 h

50% task-completion time horizon

Claude Mythos Preview (early) · harness: METR ReAct agent (Inspect)

7 Apr 2026 · Source: METR (benchmark maintainers)

About 17.4 hours. Added to the page 2026-05-08 together with a notice that measurements above 16 hours are unreliable with the current task suite. 80% horizon 185.9 minutes. Date is the release_date field in METR's data file.

Unrated

No fixed reference point (for example Elo scores or field signals).

See the full leaderboard at METR

METR Task-Completion Time Horizons: 50% task-completion time horizon over time, 6 recorded results. 0 s3.3 h6.7 h10 h13 h17 h20 hJan 2024Jan 2025Jan 2026 GPT-4 (0314): 4 min (14 Mar 2023) Claude 3.7 Sonnet: 1 h (24 Feb 2025) GPT-5: 3.4 h (7 Aug 2025) Claude Opus 4.5: 4.9 h (24 Nov 2025) Claude Opus 4.6: 12 h (5 Feb 2026) Claude Mythos Preview (early): 17 h (7 Apr 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/metr-time-horizons/"><img src="https://canagentswork.com/og/benchmarks-metr-time-horizons.png" width="600" height="315" alt="METR Task-Completion Time Horizons: the best result is 17 h (Claude Mythos Preview (early), 7 Apr 2026)." loading="lazy"></a>

Markdown:

[![METR Task-Completion Time Horizons: the best result is 17 h (Claude Mythos Preview (early), 7 Apr 2026).](https://canagentswork.com/og/benchmarks-metr-time-horizons.png)](https://canagentswork.com/benchmarks/metr-time-horizons/)

What it measures

How long a task (in expert human time) an agent can complete with 50% reliability. The Time Horizon 1.1 suite has 228 tasks from HCAST, RE-Bench, and short SWAA tasks, with 31 tasks estimated at 8 hours or more for humans. Human times come from contracted professionals with about 5 years of experience, given the same instructions and tools as the agents. Each model runs with a scaffold that METR chooses after a small elicitation phase; runs are checked for reward hacking. METR also reports an 80% horizon, which is several times shorter.

50% task-completion time horizon: Human task length at which the fitted logistic curve predicts a 50% success rate for the agent. Time Horizon 1.1 suite unless noted. Higher is better.

No fixed reference point, so we do not rate its status. Open-ended scale in human minutes. METR says measurements above 16 hours (960 minutes) are unreliable with the current suite.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli, repo
Grading
automated-tests, state-check
Tasks
228
Human reference
Skilled contractors (software, ML, cybersecurity; about 5 years of experience) attempt each task; the geometric mean of successful completion times sets the task length. Only 5 of the 31 tasks of 8 or more hours have measured human baselines; the rest use estimates.
Contamination
Most tasks are private. SWAA short tasks and some HCAST tasks are described publicly. METR screens runs for reward hacking by keyword search, LLM flags, and human review.
Reuse
Horizon estimates and run data are public in eval-analysis-public (no license file seen). Many tasks are private. (cite-only)

Limits to keep in mind

  • Measurements above 16 hours are unreliable with the current task suite (notice added 2026-05-08). The top model, Claude Mythos Preview (early), is at 1,044.8 minutes, above that line, with a 95% CI of 508.9 to 3,304.3 minutes. Source
  • Tasks are self-contained and well specified, so the horizon is closer to what a low-context new hire or contractor could do, not a professional with full context. METR says agents do worse on messier tasks and when graded holistically. Source
  • Task composition changes the trend. Moving from TH1 to TH1.1 changed the post-2023 doubling time from 165 to 131 days and moved individual estimates by up to 57%. A regularization fix on 2026-03-03 changed estimates again (Claude Opus 4.5 went from 320 to 293 minutes). Source
  • Coverage is not complete. As of 2026-09-23 METR listed Claude Opus 4.7, Grok 4.3, and GPT-5.5 as recent models without time horizons. Evaluations take 1 to 2 weeks or more. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

System50% task-completion time horizonDateSource
Claude Mythos Preview (early) · frontier
harness: METR ReAct agent (Inspect)
17 h7 Apr 2026METR · primary
Claude Opus 4.6
harness: METR ReAct agent (Inspect)
12 h5 Feb 2026METR · primary
Claude Opus 4.5
harness: METR ReAct agent (Inspect)
4.9 h24 Nov 2025METR · primary
GPT-5
harness: Triframe (Inspect)
3.4 h7 Aug 2025METR · primary
Claude 3.7 Sonnet
harness: METR ReAct agent (Inspect)
1 h24 Feb 2025METR · primary
GPT-4 (0314)
harness: modular-public
4 min14 Mar 2023METR · primary

Where sources disagree

METR's Time Horizon 1.1 blog post (2026-01-29) gives Claude Opus 4.5 a 50% horizon of 320 minutes. METR's live data file gives 293.0 minutes. Both are METR sources. The results page lists a 2026-03-03 update that "corrected a regularization mistake that affected our measurements", which explains the change. GPT-5 moved the same way (214 to 203 minutes).

  • 293metr.org (primary) · Live data file linked from metr.org/time-horizons, retrieved 2026-09-23. CI 161.7 to 623.7.
  • 320metr.org (primary) · TH1.1 launch post, table "Changes to Model Horizon Estimates". CI 170 to 729.

We show 293. Status: resolved.

Timeline

  • 8 May 2026 — METR reports a 17-hour time horizon and warns the suite is near its ceiling. Source
  • 29 Jan 2026 — METR releases Time Horizon 1.1 with 228 tasks. Source
  • 18 Mar 2025 — METR introduces the 50% task-completion time horizon. Source

Where it sits in the atlas

Software engineeringML and research engineering (partial)Security (partial)

Long-horizon autonomy

Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.

How to cite

Credit the original work first: METR Task-Completion Time Horizons by METR (https://arxiv.org/abs/2503.14499).

Then, if you used this page:

Can Agents Work. "METR Task-Completion Time Horizons: frontier results and sources." https://canagentswork.com/benchmarks/metr-time-horizons/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-metr-time-horizons,
  title        = {{METR Task-Completion Time Horizons: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/metr-time-horizons/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}