Benchmarks / OSWorld 2.0

OSWorld 2.0

Built by HKU and Snorkel AI · Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Tianbao Xie, et al. · released 26 Jun 2026

108 long-horizon computer-use workflows on a real desktop with 31 self-hosted websites, spanning research, creative production, engineering, personal services, administration, business and finance, and healthcare. A skilled user needs a median of about 1.6 hours per task. The successor to OSWorld-Verified from the XLANG Lab, released June 2026 with bug-fix releases since.

Frontier

44.3%

Binary completion rate (full set, 500 steps)

Claude Opus 5 (max, batch tool)

17 Sep 2026 · Source: HKU (benchmark maintainers)

Release v2.1 (2026-09-16), full set of 108 tasks, 500 steps, reasoning max with batch tool; official run. Partial score 77.67%. On the 82-task offline subset the same run scores 48.65% binary. Only Claude Opus 5 rows exist on v2.1. Date is the data file's updatedAt.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at HKU

OSWorld 2.0: Binary completion rate (full set, 500 steps) over time, 3 recorded results. 0%20%40%60%80%100%Jul 2026Aug 2026Sep 2026Oct 2026 Claude Opus 4.8 (max, batched tool): 20.6% (26 Jun 2026) Claude Opus 5 (max, batch tool): 31.4% (3 Sep 2026) Claude Opus 5 (max, batch tool): 44.3% (17 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/osworld-2/"><img src="https://canagentswork.com/og/benchmarks-osworld-2.png" width="600" height="315" alt="OSWorld 2.0: the best result is 44.3% (Claude Opus 5 (max, batch tool), 17 Sep 2026)." loading="lazy"></a>

Markdown:

[![OSWorld 2.0: the best result is 44.3% (Claude Opus 5 (max, batch tool), 17 Sep 2026).](https://canagentswork.com/og/benchmarks-osworld-2.png)](https://canagentswork.com/benchmarks/osworld-2/)

What it measures

Share of the 108 workflows that an agent completes end to end (binary completion), judged by state checks with an average of 27.25 scoring checkpoints per task. The agent sees the screen and acts with mouse, keyboard, and tools for up to 500 steps. Tasks mix documents, email, chat, legacy web portals, CAD, media, and medical software, and include information that arrives mid-task and cases where the agent should ask the simulated user. The site also reports a partial score (checkpoint credit) and an 82-task offline subset that runs without internet.

Binary completion rate (full set, 500 steps): Share of the 108 tasks whose final state passes every check, at a 500-step budget, on the current benchmark release. The site's default metric ("Binary Accuracy"). Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Project level: whole projects judged by an acceptance standard
Environment
computer-use, browser, documents
Grading
state-check
Tasks
108
Human reference
Skilled users take a median of about 1.6 hours per task, and 69.6% of tasks take more than one hour. No human success rate is published.
Contamination
The maintainers keep task classes and assets behind gated Hugging Face access so that agents cannot find task answers, setup logic, or evaluator details online during a run. Websites are self-hosted mocks. 82 of the 108 tasks run without internet; the rest need live sites.
Reuse
The task classes are a gated Hugging Face dataset (xlangai/osworld_v2_tasks) whose card says Apache-2.0; the task assets (xlangai/osworld_v2_assets_gated) are gated with no license on the card. The code repository is Apache-2.0. (cite-only)

Limits to keep in mind

  • Three releases in three months (v2026.06.24, v2026.08.08 on 2026-08-08, v2.1 on 2026-09-16) changed task files, assets, and evaluation workflows. Scores are listed per release, and the frontier rose with each release, so gains mix model progress with task fixes. Source
  • Verified entries require the maintainers to run the agent on their side, so the board shows few models (Claude Opus 5 only on v2.1). Lab-reported numbers use other scopes: Anthropic reports Opus 5.5 at 81.8% "partial", which is the checkpoint score, not binary completion. Source
  • Results on the 82-task offline subset and partial scores are separate scopes and run higher than full-set binary completion (Opus 5 max: 48.65% offline binary, 77.67% full-set partial). Source
  • The v2026.08.08 rows average 7 runs for Claude Opus 5 and 2 runs for GPT-5.6 Sol; the site states no run count or confidence interval for the v2.1 rows. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemBinary completion rate (full set, 500 steps)DateSource
Claude Opus 5 (max, batch tool) · frontier44.3%17 Sep 2026HKU · primary
Claude Opus 5 (max, batch tool)31.4%3 Sep 2026HKU · primary
Claude Opus 4.8 (max, batched tool)20.6%26 Jun 2026HKU · primary

Timeline

  • 16 Sep 2026 — OSWorld 2.1 bug-fix release; Claude Opus 5 completes 44.33% of the full set. Source
  • 8 Aug 2026 — OSWorld 2.0 release v2026.08.08 updates task files, assets, and mocked websites. Source
  • 26 Jun 2026 — OSWorld 2.0 launches with 108 long-horizon computer-use workflows; best agent completes 20.6%. Source

Where it sits in the atlas

Work ladder: Direct evidence for Office and administrative support. On the work ladder it counts as 44.3% of projects completed. How the ladder works

Work it measures (O*NET work activities): Operate computer systems or computerized equipment; Process digital or online data; Gather information from physical or electronic sources; Perform administrative or clerical activities; Communicate with others about operational plans or activities.

Office and administrative supportDesign, media, and writing (partial)Science and research (partial)Architecture and engineering (partial)Finance and accounting (partial)Healthcare (partial)

Computer useEnterprise workflows

Last checked 24 Sep 2026 against 9 primary sources. See an error? Tell us.

How to cite

Credit the original work first: OSWorld 2.0 by HKU and Snorkel AI (https://arxiv.org/abs/2606.29537).

Then, if you used this page:

Can Agents Work. "OSWorld 2.0: frontier results and sources." https://canagentswork.com/benchmarks/osworld-2/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-osworld-2,
  title        = {{OSWorld 2.0: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/osworld-2/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}