Can agents work?
This atlas maps 42 agent benchmarks from 37 institutions onto 997 US occupations. Color shows what the evidence says today. Hatched areas have no benchmark at the chosen level yet. Area is US employment.
The Work Atlas
- PhysicalUnexplored(Physical work, task level)
- HealthPartial(Healthcare, task level)
- ManagementPartial(Management and business operations, task level)
- OfficeStrong(Office and administrative support, task level)
- SalesPartial(Sales and marketing, task level)
- EducationUnexplored(Education and social services, task level)
- FinancePartial(Finance and accounting, task level)
- EngineeringUnexplored(Architecture and engineering, task level)
- SupportPartial(Customer support, task level)
- DevOps & ITPartial(DevOps, SRE, and IT operations, task level)
- SoftwareStrong(Software engineering, task level)
- Design & mediaPartial(Design, media, and writing, task level)
- SciencePartial(Science and research, task level)
- LegalUnexplored(Legal, task level)
- DataPartial(Data and analytics, task level)
- SecurityPartial(Security, task level)
- ML researchPartial(ML and research engineering, task level)
- PhysicalUnexplored(Physical work, task level)
- HealthPartial(Healthcare, task level)
- ManagementPartial(Management and business operations, task level)
- OfficeStrong(Office and administrative support, task level)
- SalesPartial(Sales and marketing, task level)
- EducationUnexplored(Education and social services, task level)
- FinancePartial(Finance and accounting, task level)
- EngineeringUnexplored(Architecture and engineering, task level)
- SupportPartial(Customer support, task level)
- DevOps & ITPartial(DevOps, SRE, and IT operations, task level)
- SoftwareStrong(Software engineering, task level)
- Design & mediaPartial(Design, media, and writing, task level)
- SciencePartial(Science and research, task level)
- LegalUnexplored(Legal, task level)
- DataPartial(Data and analytics, task level)
- SecurityPartial(Security, task level)
- ML researchPartial(ML and research engineering, task level)
- PhysicalUnexplored(Physical work, project level)
- HealthPartial(Healthcare, project level)
- ManagementPartial(Management and business operations, project level)
- OfficePartial(Office and administrative support, project level)
- SalesPartial(Sales and marketing, project level)
- EducationPartial(Education and social services, project level)
- FinancePartial(Finance and accounting, project level)
- EngineeringNot yet(Architecture and engineering, project level)
- SupportPartial(Customer support, project level)
- DevOps & ITUnexplored(DevOps, SRE, and IT operations, project level)
- SoftwareNot yet(Software engineering, project level)
- Design & mediaNot yet(Design, media, and writing, project level)
- ScienceNot yet(Science and research, project level)
- LegalPartial(Legal, project level)
- DataPartial(Data and analytics, project level)
- SecurityUnexplored(Security, project level)
- ML researchPartial(ML and research engineering, project level)
- PhysicalUnexplored(Physical work, project level)
- HealthPartial(Healthcare, project level)
- ManagementPartial(Management and business operations, project level)
- OfficePartial(Office and administrative support, project level)
- SalesPartial(Sales and marketing, project level)
- EducationPartial(Education and social services, project level)
- FinancePartial(Finance and accounting, project level)
- EngineeringNot yet(Architecture and engineering, project level)
- SupportPartial(Customer support, project level)
- DevOps & ITUnexplored(DevOps, SRE, and IT operations, project level)
- SoftwareNot yet(Software engineering, project level)
- Design & mediaNot yet(Design, media, and writing, project level)
- ScienceNot yet(Science and research, project level)
- LegalPartial(Legal, project level)
- DataPartial(Data and analytics, project level)
- SecurityUnexplored(Security, project level)
- ML researchPartial(ML and research engineering, project level)
- StrongAgents succeed on most benchmark tasks. Limits remain in reliability, scope, or cost.
- PartialAgents succeed on some scoped tasks. Reliable or end-to-end work is not shown.
- Not yetThe best agents fail most tasks on the best available benchmarks.
- UnexploredNo benchmark covers this work yet.
Area: US employment, BLS OEWS May 2025 (155M jobs). Color: our reading of the benchmarks mapped to each family, with sources on each family page. Occupations: O*NET 31.0.
Full-size atlas images (PNG, 1600 × 1200, with legend and sources): task level · project level.
How to cite
Credit the original work first: the benchmarks behind each job family, listed on each family page and at /sources/.
Then, if you used this page:
Can Agents Work. "The Work Atlas: AI agent benchmarks mapped onto US jobs." https://canagentswork.com/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-work-atlas,
title = {{The Work Atlas: AI agent benchmarks mapped onto US jobs}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}Can agents do my job?
Job families
| Family | US jobs | Task level | Project level | Benchmarks |
|---|---|---|---|---|
| Physical work | 60M | Unexplored | Unexplored | 0 |
| Healthcare | 18M | Partial thin evidence | Partial | 3 |
| Management and business operations | 18M | Partial thin evidence | Partial | 5 |
| Office and administrative support | 15M | Strong | Partial | 8 |
| Sales and marketing | 14M | Partial thin evidence | Partial | 3 |
| Education and social services | 12M | Unexplored | Partial thin evidence | 1 |
| Finance and accounting | 3.1M | Partial | Partial | 5 |
| Architecture and engineering | 2.6M | Unexplored | Not yet | 3 |
| Customer support | 2.6M | Partial | Partial thin evidence | 3 |
| DevOps, SRE, and IT operations | 2.4M | Partial | Unexplored | 8 |
| Software engineering | 2.2M | Strong | Not yet | 20 |
| Design, media, and writing | 2.0M | Partial | Not yet | 4 |
| Science and research | 1.5M | Partial thin evidence | Not yet thin evidence | 2 |
| Legal | 1.3M | Unexplored | Partial | 3 |
| Data and analytics | 500K | Partial | Partial | 5 |
| Security | 191K | Partial | Unexplored | 5 |
| ML and research engineering | 37K | Partial | Partial thin evidence | 6 |
Latest on the frontier
- 22 Sep 2026Scale AI releases SWE-Bench Pro V2 with 642 validated tasks
- 8 Sep 2026APEX-Agents 1.1 stops rewarding hedged answers
- 7 Sep 2026CyberGym leader reaches 98.5% reproduction rate
- 7 Sep 2026GPT-6 Astra takes the Vending-Bench 2 lead at about $15,500
- Sep 2026GDPval-AA v2.1 re-anchors the Elo scale
- 25 Jul 2026First verified OSWorld score above 90%
Recent frontier results
- 1846 EloClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) · 23 Sep 2026Unrated
The atlas stands on their work
Every benchmark here was built by someone else. We summarize, map, and link. Start with their work.
aavetisAiderAndon LabsAnthropicArtificial AnalysisBoxCAISCarnegie Mellon UniversityCisco ResearchConcordia UniversityHarveyHKUHugging FaceIBM ResearchIIScLaude InstituteMartianMercorMetaMETRMicrosoft ResearchNebiusOpenAIPrinceton UniversityQueen's UniversityRenmin University of ChinaSalesforce AI ResearchScale AIServiceNowSierraStanford UniversityUC BerkeleyUIUCUniversity of MichiganUniversity of TorontoUniversity of WaterlooVals AI