Benchmarks / ITBench

ITBench

Built by IBM Research and UIUC · released 7 Feb 2025

IT automation scenarios that run in live Kubernetes environments. Agents must diagnose and repair incidents (SRE), assess and enforce compliance (CISO), and find cost problems (FinOps). IBM Research hosts the environments and a leaderboard.

Frontier

25%

SRE incidents resolved

ITBench-SRE-Agent-GPT-4o · harness: ITBench-SRE-Agent

2 May 2025 · Source: IBM Research (benchmark maintainers)

Single-trial table, 16 trials across incidents. The multi-trial table (162 trials) shows 24.79% for the same agent. Leaderboard date read as 2 May 2025.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at IBM Research

ITBench: SRE incidents resolved over time, 3 recorded results. 0%20%40%60%80%100%May 2025Jun 2025Jul 2025Aug 2025 ITBench-SRE-Agent-LLama-3-3-70B: 12.5% (2 May 2025) ITBench-SRE-Agent-GPT-4o: 25% (2 May 2025) Agents powered by state-of-the-art models (paper, ICML 2025): 11.4% (Jul 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/itbench/"><img src="https://canagentswork.com/og/benchmarks-itbench.png" width="600" height="315" alt="ITBench: the best result is 25% (ITBench-SRE-Agent-GPT-4o, 2 May 2025)." loading="lazy"></a>

Markdown:

[![ITBench: the best result is 25% (ITBench-SRE-Agent-GPT-4o, 2 May 2025).](https://canagentswork.com/og/benchmarks-itbench.png)](https://canagentswork.com/benchmarks/itbench/)

What it measures

For SRE scenarios, whether the agent localizes the fault and repairs the incident so that the alert clears. The headline number here is the share of SRE incidents resolved. The CISO track scores compliance assessments and policy generation (1.0 is perfect). The FinOps track scores cost optimization and anomaly detection. The ICML 2025 paper covers 102 scenarios; the public repo ships a smaller open-source subset.

SRE incidents resolved: Share of SRE incidents that the agent repaired, as reported on the ITBench SRE leaderboard. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
live-system, cli
Grading
state-check, automated-tests
Tasks
102
Reuse
Apache-2.0 (repository LICENSE file). Scenario tooling and sample scenarios are open source; hosted environments require registration. (open-apache)

Limits to keep in mind

  • The SRE leaderboard was last updated in May 2025 and lists only IBM's own reference agent with three models. Newer models have not been posted there. Source
  • The paper's headline (11.4% of SRE scenarios resolved) and the leaderboard's best entry (25.0%) use different scenario sets and agents, so they are not directly comparable. Source
  • SREGym, a later benchmark from some of the same authors, reports that problems ported from ITBench and AIOpsLab are now easy for strong agents (mitigation above 80%). Source
  • The arXiv v1 abstract (94 scenarios, 13.8% SRE) and the ICML version (102 scenarios, 11.4% SRE) report different counts and rates. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemSRE incidents resolvedDateSource
Agents powered by state-of-the-art models (paper, ICML 2025)
harness: ITBench reference agents
11.4%Jul 2025IBM Research · primary
ITBench-SRE-Agent-LLama-3-3-70B
harness: ITBench-SRE-Agent
12.5%2 May 2025IBM Research · primary
ITBench-SRE-Agent-GPT-4o · frontier
harness: ITBench-SRE-Agent
25%2 May 2025IBM Research · primary

Where sources disagree

The ICML 2025 paper says agents resolve 11.4% of SRE scenarios. The arXiv v1 abstract says 13.8%. The official SRE leaderboard shows 25.0% for IBM's GPT-4o reference agent. The three numbers come from different scenario sets and dates. We show the leaderboard value because it is the maintainers' most recent published result.

  • 11.4proceedings.mlr.press (primary) · ICML 2025 camera-ready abstract, 102 scenarios.
  • 13.8arxiv.org (primary) · arXiv v1 abstract (2025-02-07), 94 scenarios.
  • 25github.com (primary) · SRE leaderboard, ITBench-SRE-Agent-GPT-4o, 16 trials, updated 2 May 2025. Multi-trial table shows 24.79%.

We show 25. Status: open.

Timeline

  • 7 Feb 2025 — IBM Research releases ITBench. Source

Where it sits in the atlas

DevOps, SRE, and IT operationsSecurity (partial)

Incident responseCompliance operationsFinOps

Last checked 23 Sep 2026 against 5 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: ITBench by IBM Research and UIUC (https://proceedings.mlr.press/v267/jha25a.html).

Then, if you used this page:

Can Agents Work. "ITBench: frontier results and sources." https://canagentswork.com/benchmarks/itbench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-itbench,
  title        = {{ITBench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/itbench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}