Benchmarks / EnterpriseOps-Gym
EnterpriseOps-Gym
Built by ServiceNow, Mila and Université de Montréal · Shiva Krishna Reddy Malay, Shravan Nayak, Jishnu Sethumadhavan Nair, Aman Tiwari, et al. · released 13 Mar 2026
A containerized enterprise simulation from ServiceNow AI Research. An agent works through MCP tools against live databases in eight domains: Calendar, Customer Service Management, Drive, Email, HR, IT Service Management, Teams, and Hybrid tasks that span several systems. Expert-written SQL checks grade the final state, not the actions.
Frontier
45.9%
Task success rate
Claude Opus 4.6 · harness: ReAct (EnterpriseOps-Gym)
16 Mar 2026 · Source: ServiceNow (benchmark maintainers)
README leaderboard, full benchmark, oracle mode. By domain: Teams 52.0, CSM 45.1, Email 57.7, ITSM 33.3, Calendar 43.3, HR 45.1, Drive 57.1, Hybrid 34.0. The table is undated; it appears in the repository's first commit (2026-03-16). Artificial Analysis's separate run of the public split scores Claude Fable 5 at 51.1%, which is not comparable.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Task success rate: a task passes only when every verification condition on the final database state holds (goal reached, data intact, policy followed, no side effects). 1,150 expert-curated tasks, 512 tools, and 164 tables; expert trajectories average 9.15 steps (up to 34). The headline setting is oracle tool mode, where the agent gets the tools the task needs. 30 tasks are infeasible and test whether the agent refuses cleanly.
Task success rate: Share of tasks where all SQL verifiers pass, in oracle tool mode on the full benchmark. The README leaderboard reports the mean of the eight domain rates. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- api-tools, simulated-workplace
- Grading
- state-check
- Tasks
- 1,150
- Contamination
- 60% of the tasks are public. The README leaderboard reports the full benchmark and, separately, the public split.
- Reuse
- Apache 2.0 on the Hugging Face dataset card and the repository. The public split holds 60% of the tasks (649 rows in the oracle config); the rest is private. (open-apache)
Limits to keep in mind
- Three numbers for the best model appear at the source: the README intro says 34.1%, the paper says 37.4% (Claude Opus 4.5), and the README leaderboard says 45.9% (Claude Opus 4.6). The paper averages over tasks; the README board averages the eight domain rates. See the conflict file. Source
- The maintainers' site now names the Artificial Analysis board the official leaderboard. That board runs the public dataset in AA's own Stirrup harness with 3 repeats, and AA says its numbers are not directly comparable with the paper. Its top score is 51.1% (Claude Fable 5, Opus 4.8 fallback), undated. Source
- Oracle tool mode gives the agent the right tools. The paper reports that adding 5 to 15 distractor tools changed Claude Sonnet 4.5's score by about one point, so the effect of tool retrieval is small but the headline is still the easiest setting. Source
- Agents refuse infeasible tasks cleanly only 53.9% of the time at best, so side effects on policy-violating requests are common. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Task success rate | Date | Source |
|---|---|---|---|
| Claude Opus 4.5 harness: ReAct (EnterpriseOps-Gym) | 39.4% | 16 Mar 2026 | ServiceNow · primary |
| Claude Opus 4.6 · frontier harness: ReAct (EnterpriseOps-Gym) | 45.9% | 16 Mar 2026 | ServiceNow · primary |
Where sources disagree
Three "best model" numbers appear at the source. The paper says Claude Opus 4.5 reaches 37.4%. The README leaderboard shows 39.4% for the same model, with the same eight domain scores. The README intro says "best model achieves only 34.1%", which equals the board's row for Claude Sonnet 4.5, while the board's top row is Claude Opus 4.6 at 45.9%.
- 37.4 — arxiv.org (primary) · Paper abstract and Table 2 (2026-03-13), Claude Opus 4.5, oracle tool mode.
- 39.4 — github.com (primary) · README leaderboard "Avg" for Claude Opus 4.5, full benchmark, oracle mode.
- 34.1 — github.com (primary) · README intro line "Best model achieves only 34.1% success rate". Matches the board row for Claude Sonnet 4.5.
We show 39.4. Status: resolved.
Timeline
- 13 Mar 2026 — ServiceNow releases EnterpriseOps-Gym; best model passes 37.4% of enterprise tasks. Source
Where it sits in the atlas
Work ladder: Direct evidence for Office and administrative support. On the work ladder it counts as 45.9% of tasks completed. How the ladder works
Work it measures (O*NET work activities): Perform administrative or clerical activities; Schedule appointments; Maintain operational records; Communicate with others about operational plans or activities; Respond to customer problems or inquiries.
Office and administrative supportCustomer support (partial)DevOps, SRE, and IT operations (partial)Management and business operations (partial)
Go to the source
- Website enterpriseops-gym.github.io
- Paper arxiv.org
- Full leaderboard github.com
- Code github.com
- Dataset huggingface.co
Last checked 24 Sep 2026 against 8 primary sources. See an error? Tell us.
How to cite
Credit the original work first: EnterpriseOps-Gym by ServiceNow, Mila, and Université de Montréal (https://arxiv.org/abs/2603.13594).
Then, if you used this page:
Can Agents Work. "EnterpriseOps-Gym: frontier results and sources." https://canagentswork.com/benchmarks/enterpriseops-gym/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-enterpriseops-gym,
title = {{EnterpriseOps-Gym: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/enterpriseops-gym/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}