Benchmarks / HealthAgentBench
HealthAgentBench
Built by Microsoft Research · Qianchu Liu, Sheng Zhang, Guanghui Qin, Jeya Maria Jose Valanarasu, et al. · released 1 Jul 2026
54 terminal tasks from Microsoft Research in seven healthcare categories: correct a chest X-ray report, mark tumor tiles on a pathology slide, classify abnormalities in a CT volume, match a patient to clinical trials, audit an EHR table for injected errors, build a risk model over EHR timelines, and customize an EHR ETL pipeline. A coding agent gets minimal instructions and real clinical data, and a hidden gold label decides success.
Frontier
55%
Task success rate
Claude Code (Opus 5) · harness: Claude Code
27 Jul 2026 · Source: Microsoft Research (benchmark maintainers)
89 of 162 trials (0.5494 in results.json); the site shows 55%. Cost $3.3 per task. Added 2026-07-27 per the README. Per category: trial matching 78%, EHR event modelling 72%, EHR audit 71%, pathology 57%, X-ray 33%, CT 27%, ETL 100%.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Pooled task success rate: the share of 162 trials (3 attempts on each of 54 tasks) that meet the task's binary success criterion, for example zero clinically significant errors in the corrected report, tile F1 of at least 0.9, all CT labels correct, full recall on the audit and trial-matching tasks, or matching a human-engineered baseline on event prediction. The maintainers run each agent CLI (Codex, Claude Code, Copilot CLI) at xhigh reasoning effort with web browsing disabled, in Harbor, and report mean cost and wall-clock time per task.
Task success rate: Share of the 162 trials (3 attempts × 54 tasks) that pass the task's success criterion. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli
- Grading
- automated-tests, llm-judge
- Tasks
- 54
- Contamination
- Source datasets are public research sets, so their labels exist online. The maintainers disable web search and fetch in the agent CLIs and keep gold labels out of the repository.
- Reuse
- Code and task definitions are MIT (repository LICENSE). Gold labels are not in the repository; the run fetches them. Three categories need gated source data with per-user credentials: MIMIC-CXR (PhysioNet), CT-RATE (Hugging Face, OpenRAIL gated), and EHRSHOT (Stanford Redivis). The other categories use public data (CAMELYON16, MIMIC-IV demo, TREC-CT 2021). (cite-only)
Limits to keep in mind
- Small suite: 54 tasks, and the EHR Format Conversion category has one task. The leaderboard shows whole percents; the site's results.json holds the exact fractions. Source
- Three of seven categories need credentials for gated datasets (MIMIC-CXR, CT-RATE, EHRSHOT), so a full run is not open to everyone. Source
- X-ray report correction is scored by an LLM judge (CheXprompt with GPT-5.4 by default), so one category depends on the judge model. Source
- Rows mix a model and an agent CLI: GPT-5.5 scores 42% under Codex and 35% under Copilot CLI. All rows are the maintainers' runs of OpenAI and Anthropic models. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Task success rate | Date | Source |
|---|---|---|---|
| Codex (GPT-5.6-sol) harness: Codex | 45% | 27 Jul 2026 | Microsoft Research · primary |
| Claude Code (Opus 5) · frontier harness: Claude Code | 55% | 27 Jul 2026 | Microsoft Research · primary |
| Codex (GPT 5.5) harness: Codex | 42% | 1 Jul 2026 | Microsoft Research · primary |
Timeline
Where it sits in the atlas
Work ladder: Direct evidence for Healthcare. On the work ladder it counts as 55% of tasks completed. How the ladder works
Work it measures (O*NET work activities): Analyze health or medical data; Prepare health or medical documents; Maintain health or medical records.
HealthcareScience and research (partial)Data and analytics (partial)
Medical image readingClinical records workData engineering and SQL
Go to the source
- Website microsoft.github.io
- Paper arxiv.org
- Full leaderboard microsoft.github.io
- Code github.com
Last checked 24 Sep 2026 against 4 primary sources. See an error? Tell us.
How to cite
Credit the original work first: HealthAgentBench by Microsoft Research (https://arxiv.org/abs/2606.31179).
Then, if you used this page:
Can Agents Work. "HealthAgentBench: frontier results and sources." https://canagentswork.com/benchmarks/healthagentbench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-healthagentbench,
title = {{HealthAgentBench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/healthagentbench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}