Benchmarks / HealthAgentBench

HealthAgentBench

Built by Microsoft Research · Qianchu Liu, Sheng Zhang, Guanghui Qin, Jeya Maria Jose Valanarasu, et al. · released 1 Jul 2026

54 terminal tasks from Microsoft Research in seven healthcare categories: correct a chest X-ray report, mark tumor tiles on a pathology slide, classify abnormalities in a CT volume, match a patient to clinical trials, audit an EHR table for injected errors, build a risk model over EHR timelines, and customize an EHR ETL pipeline. A coding agent gets minimal instructions and real clinical data, and a hidden gold label decides success.

Frontier

55%

Task success rate

Claude Code (Opus 5) · harness: Claude Code

27 Jul 2026 · Source: Microsoft Research (benchmark maintainers)

89 of 162 trials (0.5494 in results.json); the site shows 55%. Cost $3.3 per task. Added 2026-07-27 per the README. Per category: trial matching 78%, EHR event modelling 72%, EHR audit 71%, pathology 57%, X-ray 33%, CT 27%, ETL 100%.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at Microsoft Research

HealthAgentBench: Task success rate over time, 3 recorded results. 0%20%40%60%80%100%Jul 2026Aug 2026 Codex (GPT 5.5): 42% (1 Jul 2026) Codex (GPT-5.6-sol): 45% (27 Jul 2026) Claude Code (Opus 5): 55% (27 Jul 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/healthagentbench/"><img src="https://canagentswork.com/og/benchmarks-healthagentbench.png" width="600" height="315" alt="HealthAgentBench: the best result is 55% (Claude Code (Opus 5), 27 Jul 2026)." loading="lazy"></a>

Markdown:

[![HealthAgentBench: the best result is 55% (Claude Code (Opus 5), 27 Jul 2026).](https://canagentswork.com/og/benchmarks-healthagentbench.png)](https://canagentswork.com/benchmarks/healthagentbench/)

What it measures

Pooled task success rate: the share of 162 trials (3 attempts on each of 54 tasks) that meet the task's binary success criterion, for example zero clinically significant errors in the corrected report, tile F1 of at least 0.9, all CT labels correct, full recall on the audit and trial-matching tasks, or matching a human-engineered baseline on event prediction. The maintainers run each agent CLI (Codex, Claude Code, Copilot CLI) at xhigh reasoning effort with web browsing disabled, in Harbor, and report mean cost and wall-clock time per task.

Task success rate: Share of the 162 trials (3 attempts × 54 tasks) that pass the task's success criterion. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli
Grading
automated-tests, llm-judge
Tasks
54
Contamination
Source datasets are public research sets, so their labels exist online. The maintainers disable web search and fetch in the agent CLIs and keep gold labels out of the repository.
Reuse
Code and task definitions are MIT (repository LICENSE). Gold labels are not in the repository; the run fetches them. Three categories need gated source data with per-user credentials: MIMIC-CXR (PhysioNet), CT-RATE (Hugging Face, OpenRAIL gated), and EHRSHOT (Stanford Redivis). The other categories use public data (CAMELYON16, MIMIC-IV demo, TREC-CT 2021). (cite-only)

Limits to keep in mind

  • Small suite: 54 tasks, and the EHR Format Conversion category has one task. The leaderboard shows whole percents; the site's results.json holds the exact fractions. Source
  • Three of seven categories need credentials for gated datasets (MIMIC-CXR, CT-RATE, EHRSHOT), so a full run is not open to everyone. Source
  • X-ray report correction is scored by an LLM judge (CheXprompt with GPT-5.4 by default), so one category depends on the judge model. Source
  • Rows mix a model and an agent CLI: GPT-5.5 scores 42% under Codex and 35% under Copilot CLI. All rows are the maintainers' runs of OpenAI and Anthropic models. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemTask success rateDateSource
Codex (GPT-5.6-sol)
harness: Codex
45%27 Jul 2026Microsoft Research · primary
Claude Code (Opus 5) · frontier
harness: Claude Code
55%27 Jul 2026Microsoft Research · primary
Codex (GPT 5.5)
harness: Codex
42%1 Jul 2026Microsoft Research · primary

Timeline

  • 27 Jul 2026 — Claude Code with Opus 5 reaches 55% on HealthAgentBench. Source
  • 1 Jul 2026 — Microsoft Research releases HealthAgentBench; Codex with GPT-5.5 succeeds on 42% of trials. Source

Where it sits in the atlas

Work ladder: Direct evidence for Healthcare. On the work ladder it counts as 55% of tasks completed. How the ladder works

Work it measures (O*NET work activities): Analyze health or medical data; Prepare health or medical documents; Maintain health or medical records.

HealthcareScience and research (partial)Data and analytics (partial)

Medical image readingClinical records workData engineering and SQL

Last checked 24 Sep 2026 against 4 primary sources. See an error? Tell us.

How to cite

Credit the original work first: HealthAgentBench by Microsoft Research (https://arxiv.org/abs/2606.31179).

Then, if you used this page:

Can Agents Work. "HealthAgentBench: frontier results and sources." https://canagentswork.com/benchmarks/healthagentbench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-healthagentbench,
  title        = {{HealthAgentBench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/healthagentbench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}