Can agents work?

This atlas maps 42 agent benchmarks from 37 institutions onto 997 US occupations. Color shows what the evidence says today. Hatched areas have no benchmark at the chosen level yet. Area is US employment.

39%of US jobs are in job families that no benchmark in this atlas covers.
42benchmarks, with 168 results that each link to a source.
37institutions built the benchmarks. See them all.
11benchmarks are saturated, solved, or retired. The frontier moves fast.

The Work Atlas

Show readings at

Area: US employment, BLS OEWS May 2025 (155M jobs). Color: our reading of the benchmarks mapped to each family, with sources on each family page. Occupations: O*NET 31.0.

Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/"><img src="https://canagentswork.com/og/home.png" width="600" height="315" alt="Can agents work? 39% of US jobs have no agent benchmark. The Work Atlas maps 42 benchmarks onto the job market." loading="lazy"></a>

Markdown:

[![Can agents work? 39% of US jobs have no agent benchmark. The Work Atlas maps 42 benchmarks onto the job market.](https://canagentswork.com/og/home.png)](https://canagentswork.com/)

Full-size atlas images (PNG, 1600 × 1200, with legend and sources): task level · project level.

How to cite

Credit the original work first: the benchmarks behind each job family, listed on each family page and at /sources/.

Then, if you used this page:

Can Agents Work. "The Work Atlas: AI agent benchmarks mapped onto US jobs." https://canagentswork.com/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-work-atlas,
  title        = {{The Work Atlas: AI agent benchmarks mapped onto US jobs}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}

Can agents do my job?

Job families

FamilyUS jobsTask levelProject levelBenchmarks
Physical work60MUnexploredUnexplored0
Healthcare18MPartial thin evidencePartial3
Management and business operations18MPartial thin evidencePartial5
Office and administrative support15MStrongPartial8
Sales and marketing14MPartial thin evidencePartial3
Education and social services12MUnexploredPartial thin evidence1
Finance and accounting3.1MPartialPartial5
Architecture and engineering2.6MUnexploredNot yet3
Customer support2.6MPartialPartial thin evidence3
DevOps, SRE, and IT operations2.4MPartialUnexplored8
Software engineering2.2MStrongNot yet20
Design, media, and writing2.0MPartialNot yet4
Science and research1.5MPartial thin evidenceNot yet thin evidence2
Legal1.3MUnexploredPartial3
Data and analytics500KPartialPartial5
Security191KPartialUnexplored5
ML and research engineering37KPartialPartial thin evidence6

Latest on the frontier

  1. 22 Sep 2026Scale AI releases SWE-Bench Pro V2 with 642 validated tasks
  2. 8 Sep 2026APEX-Agents 1.1 stops rewarding hedged answers
  3. 7 Sep 2026CyberGym leader reaches 98.5% reproduction rate
  4. 7 Sep 2026GPT-6 Astra takes the Vending-Bench 2 lead at about $15,500
  5. Sep 2026GDPval-AA v2.1 re-anchors the Elo scale
  6. 25 Jul 2026First verified OSWorld score above 90%

Full timeline · RSS

Recent frontier results

  • PR Arenaby aavetisTask level · Merge rate of ready PRs
    96%Copilot · 23 Sep 2026
    Unrated
  • GDPval-AAby Artificial AnalysisProject level · GDPval-AA Elo
    1846 EloClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) · 23 Sep 2026
    Unrated
  • APEX-Agentsby Mercor, Box and HarveyProject level · Pass@1
    73.5%Claude Opus 5.5 (max) · 23 Sep 2026
    Strong
  • GAIAby Meta and Hugging FaceTask level · Test-set accuracy
    93.7%Ops-Agentic-Search-2.0 · 22 Sep 2026
    Saturated
  • Vending-Bench 2by Andon LabsProject level · Final bank balance
    $15,515GPT-6 Astra · 7 Sep 2026
    Unrated
  • CyberGymby UC BerkeleyTask level · Success rate (Level 1)
    98.5%Creation (天工), multi-model · 7 Sep 2026
    Saturated

All benchmarks

The atlas stands on their work

Every benchmark here was built by someone else. We summarize, map, and link. Start with their work.

aavetisAiderAndon LabsAnthropicArtificial AnalysisBoxCAISCarnegie Mellon UniversityCisco ResearchConcordia UniversityHarveyHKUHugging FaceIBM ResearchIIScLaude InstituteMartianMercorMetaMETRMicrosoft ResearchNebiusOpenAIPrinceton UniversityQueen's UniversityRenmin University of ChinaSalesforce AI ResearchScale AIServiceNowSierraStanford UniversityUC BerkeleyUIUCUniversity of MichiganUniversity of TorontoUniversity of WaterlooVals AI