Jobs / Software engineering
Can agents do software engineering work?
Build, test, and maintain software, from issue fixes to whole applications.
- US jobs
- 2,150,380 (1.4% of employment)
- Typical wage
- $129K a year (employment-weighted mean of occupation medians)
- Occupations
- 7 in O*NET 31.0
- Evidence
- 20 benchmarks from 19 institutions
Task level
Strong
On fresh or held-out issue sets, the best agents resolve most tasks: 61.5% on SWE-Bench Pro (public set) and 64.5% on SWE-rebench. Limits: the SWE-Bench Pro private set is lower (51.5%), agent PRs in popular repositories are accepted less often than human PRs (65.3% vs 76.8% in AIDev), and GPT-5 cheated on 54% of impossible tasks in ImpossibleBench.
Evidence: SWE-Bench Pro, SWE-rebench, SWE-Lancer, Aider Polyglot leaderboard, Code Review Bench, Ambig-SWE, AIDev, PR Arena, ImpossibleBench, METR Task-Completion Time Horizons — by aavetis, Aider, Anthropic, Carnegie Mellon University, Martian, METR, Nebius, OpenAI, Queen's University, and Scale AI. Reading dated 23 Sep 2026.
Agents succeed on most benchmark tasks. Limits remain in reliability, scope, or cost.
Project level
Not yet
Whole projects are mostly out of reach. On the Remote Labor Index, the best agent completed 15.8% of paid freelance projects to a standard a client would accept. In CodeClash, where agents evolve their own codebases over many rounds, a human-written bot beat the best model in 150 of 150 rounds. GDPval deliverables score higher, but those tasks are one-shot and fully specified.
Evidence: Remote Labor Index, CodeClash, GDPval, GDPval-AA — by Artificial Analysis, CAIS, OpenAI, Princeton University, Scale AI, and Stanford University. Reading dated 23 Sep 2026.
The best agents fail most tasks on the best available benchmarks.
Benchmarks for this work
- 1385 EloClaude Sonnet 4.5 · 3 Nov 2025Unrated
- 79.2%Sonar Foundation Agent + Claude 4.5 Opus · 5 Dec 2025Retired
- 42.9%TTE-MatrixAgent + DeepSeek-V3.2 · 10 Nov 2025Emerging
Also relevant
These benchmarks cover part of this family's work.
- 1846 EloClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) · 23 Sep 2026Unrated
- 96.7%Genloop's Sentinel Agent v2 Pro · 1 Mar 2026Saturated
- 84.7%NexAU-AHE + GPT-5.5 · 23 Apr 2026Retired
Functions: Feature developmentShipping pull requestsAsking for clarificationIssue resolutionCode reviewCodebase evolutionLong-horizon autonomyVulnerability researchProfessional deliverablesTest integrityFreelance projectsData engineering and SQLTerminal operationsEnterprise workflows
What we do not know
- Benchmarks measure tasks, not whole jobs. A job is a bundle of tasks, judgment, and relationships.
- Capability is not adoption. A task that agents can do in a benchmark may still be done by people at work.
- Most results come from single runs. Reliability over many runs is often unknown.
Occupations in this family
Sorted by US employment (BLS OEWS May 2025). O*NET sub-occupations share the employment figure of their SOC group.
| Occupation | SOC | US jobs |
|---|---|---|
| Software Developers | 15-1252.00 | 1.7M |
| Blockchain Engineers | 15-1299.07 | (435K group) |
| Software Quality Assurance Analysts and Testers | 15-1253.00 | 187K |
| Video Game Designers | 15-1255.01 | (113K group) |
| Web and Digital Interface Designers | 15-1255.00 | 113K |
| Computer Programmers | 15-1251.00 | 92K |
| Web Developers | 15-1254.00 | 70K |
How to cite
Credit the original work first: the benchmarks listed on this page, and their institutions.
Then, if you used this page:
Can Agents Work. "Can agents do software engineering work?." https://canagentswork.com/jobs/software-engineering/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-jobs-software-engineering,
title = {{Can agents do software engineering work?}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/jobs/software-engineering/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}