Jobs / DevOps, SRE, and IT operations
Can agents do DevOps, SRE, and IT operations work?
Run systems and networks, respond to incidents, manage infrastructure, and support users.
- US jobs
- 2,382,350 (1.5% of employment)
- Typical wage
- $94K a year (employment-weighted mean of occupation medians)
- Occupations
- 11 in O*NET 31.0
- Evidence
- 8 benchmarks from 16 institutions
Task level
Partial
Mixed. On SREGym, the best agent fixes 72.2% of injected faults end to end. On ITBench-AA, no model names the root cause of more than half of Kubernetes incidents (best 47%). Infrastructure as code has only generation tests with 2024-era results (IaC-Eval, 36.7%). We found no public benchmark with recent results that tests agents making Terraform changes to live infrastructure.
Evidence: SREGym, ITBench-AA, ITBench, AIOpsLab, IaC-Eval, Terminal-Bench 2.0 — by Artificial Analysis, Cisco Research, IBM Research, IISc, Laude Institute, Microsoft Research, Stanford University, UC Berkeley, UIUC, University of Michigan, and University of Toronto. Reading dated 23 Sep 2026.
Agents succeed on some scoped tasks. Reliable or end-to-end work is not shown.
Project level
Unexplored
No benchmark in our collection tests operations work at project scale, such as a multi-day migration, a platform upgrade, or an on-call rotation.
No benchmark covers this work yet.
Benchmarks for this work
- 36.7%GPT-4 with retrieval-augmented generation · 2024Emerging
- 47%Claude Opus 4.7 (Adaptive Reasoning, Max Effort) · 27 May 2026Emerging
- 84.7%NexAU-AHE + GPT-5.5 · 23 Apr 2026Retired
Also relevant
These benchmarks cover part of this family's work.
- OSWorld-Verifiedby HKU, Salesforce AI Research, Carnegie Mellon University and University of Waterloo90.2%Intelligence-Indeed Agent · 25 Jul 2026Saturated
Functions: Incident responseTerminal operationsInfrastructure as codeCompliance operationsFinOpsComputer useSpreadsheet workFeature developmentEnterprise workflows
What we do not know
- Benchmarks measure tasks, not whole jobs. A job is a bundle of tasks, judgment, and relationships.
- Capability is not adoption. A task that agents can do in a benchmark may still be done by people at work.
- Most results come from single runs. Reliability over many runs is often unknown.
Occupations in this family
Sorted by US employment (BLS OEWS May 2025). O*NET sub-occupations share the employment figure of their SOC group.
| Occupation | SOC | US jobs |
|---|---|---|
| Computer User Support Specialists | 15-1232.00 | 717K |
| Computer Systems Analysts | 15-1211.00 | 520K |
| Health Informatics Specialists | 15-1211.01 | (520K group) |
| Computer Occupations, All Other | 15-1299.00 | 435K |
| Computer Systems Engineers/Architects | 15-1299.08 | (435K group) |
| Web Administrators | 15-1299.01 | (435K group) |
| Network and Computer Systems Administrators | 15-1244.00 | 314K |
| Computer Network Architects | 15-1241.00 | 180K |
| Telecommunications Engineering Specialists | 15-1241.01 | (180K group) |
| Computer Network Support Specialists | 15-1231.00 | 146K |
| Database Administrators | 15-1242.00 | 70K |
How to cite
Credit the original work first: the benchmarks listed on this page, and their institutions.
Then, if you used this page:
Can Agents Work. "Can agents do DevOps, SRE, and IT operations work?." https://canagentswork.com/jobs/devops-sre-it/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-jobs-devops-sre-it,
title = {{Can agents do DevOps, SRE, and IT operations work?}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/jobs/devops-sre-it/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}