Benchmarks / ITBench-AA

ITBench-AA

Built by Artificial Analysis and IBM Research · released 27 May 2026

Kubernetes incident diagnosis tasks built from IBM's ITBench and run by Artificial Analysis in a fixed agent harness. The model reads an offline incident snapshot with alerts, logs, traces, metrics, and topology, then names the root-cause Kubernetes entities.

Frontier

47%

ITBench-AA score

Claude Opus 4.7 (Adaptive Reasoning, Max Effort) · harness: Stirrup

27 May 2026 · Source: Artificial Analysis (benchmark maintainers)

Led the launch leaderboard. Most expensive at $5.38 per task.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at Artificial Analysis

ITBench-AA: ITBench-AA score over time, 4 recorded results. 0%20%40%60%80%100%May 2026May 2026May 2026May 2026Jun 2026Jun 2026 Gemini 3.1 Pro Preview: 30% (27 May 2026) GLM-5.1 (Reasoning): 40% (27 May 2026) GPT-5.5 (xhigh): 46% (27 May 2026) Claude Opus 4.7 (Adaptive Reasoning, Max Effort): 47% (27 May 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/itbench-aa/"><img src="https://canagentswork.com/og/benchmarks-itbench-aa.png" width="600" height="315" alt="ITBench-AA: the best result is 47% (Claude Opus 4.7 (Adaptive Reasoning, Max Effort), 27 May 2026)." loading="lazy"></a>

Markdown:

[![ITBench-AA: the best result is 47% (Claude Opus 4.7 (Adaptive Reasoning, Max Effort), 27 May 2026).](https://canagentswork.com/og/benchmarks-itbench-aa.png)](https://canagentswork.com/benchmarks/itbench-aa/)

What it measures

Whether the model finds the exact set of root-cause entities for a Kubernetes incident. Scoring is average precision at full recall: a repeat scores 0 if any true root cause is missed, and otherwise scores the precision of the submitted list. The headline is the mean over 59 tasks and 3 repeats. Models run in the open-source Stirrup harness with shell access and a 100-turn cap.

ITBench-AA score: Mean average precision at full recall over 59 SRE tasks x 3 repeats. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli
Grading
state-check
Tasks
59
Contamination
19 of the 59 tasks are new and held out. The other 40 are public ITBench scenarios.
Reuse
40 public tasks and 19 held-out tasks; the held-out tasks are private to Artificial Analysis. (cite-only)

Limits to keep in mind

  • Diagnosis only. Models name root-cause entities from a snapshot; they do not repair a live system. Repair is what the original ITBench SRE track scores. Source
  • Strict scoring: naming one extra entity lowers the score, and missing one true root cause gives zero for that repeat. Models that investigate longer tend to add false positives. Source
  • FinOps and CISO tasks were announced but not yet included at launch. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemITBench-AA scoreDateSource
Gemini 3.1 Pro Preview
harness: Stirrup
30%27 May 2026Artificial Analysis · primary
GLM-5.1 (Reasoning)
harness: Stirrup
40%27 May 2026Artificial Analysis · primary
GPT-5.5 (xhigh)
harness: Stirrup
46%27 May 2026Artificial Analysis · primary
Claude Opus 4.7 (Adaptive Reasoning, Max Effort) · frontier
harness: Stirrup
47%27 May 2026Artificial Analysis · primary

Timeline

  • 27 May 2026 — Artificial Analysis and IBM launch ITBench-AA. Source

Where it sits in the atlas

DevOps, SRE, and IT operations

Incident response

Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.

How to cite

Credit the original work first: ITBench-AA by Artificial Analysis and IBM Research.

Then, if you used this page:

Can Agents Work. "ITBench-AA: frontier results and sources." https://canagentswork.com/benchmarks/itbench-aa/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-itbench-aa,
  title        = {{ITBench-AA: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/itbench-aa/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}