Jobs / Office and administrative support
Can agents do office and administrative support work?
Scheduling, records, data entry, bookkeeping clerks, and other office work.
- US jobs
- 15,157,670 (9.7% of employment)
- Typical wage
- $50K a year (employment-weighted mean of occupation medians)
- Occupations
- 55 in O*NET 31.0
- Evidence
- 8 benchmarks from 10 institutions
Task level
Strong
On computer-use and office-task benchmarks, top agents pass most tasks: 90.2% on OSWorld-Verified (humans: 72.4%), 83.1% on SpreadsheetBench, and 63.3% on WorkArena L1 under the standard protocol. Limits: many top entries are self-submitted and not re-run by the maintainers, and on TheAgentCompany's multi-step workplace tasks the best agent resolves 42.9%.
Evidence: OSWorld-Verified, SpreadsheetBench, WorkArena, GAIA, BrowseComp, TheAgentCompany — by Carnegie Mellon University, HKU, Hugging Face, Meta, OpenAI, Renmin University of China, Salesforce AI Research, ServiceNow, and University of Waterloo. Reading dated 23 Sep 2026.
Agents succeed on most benchmark tasks. Limits remain in reliability, scope, or cost.
Project level
Partial
Only adjacent evidence. OpenAI reports that GPT-5.5's GDPval deliverables won or tied against professionals' work in 84.9% of comparisons, but GDPval tasks are one-shot and fully specified. No benchmark in our collection tests multi-day administrative work, such as running a calendar or a purchasing process.
Evidence: GDPval, GDPval-AA — by Artificial Analysis and OpenAI. Reading dated 23 Sep 2026.
Agents succeed on some scoped tasks. Reliable or end-to-end work is not shown.
Benchmarks for this work
- OSWorld-Verifiedby HKU, Salesforce AI Research, Carnegie Mellon University and University of Waterloo90.2%Intelligence-Indeed Agent · 25 Jul 2026Saturated
Also relevant
These benchmarks cover part of this family's work.
- 1846 EloClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) · 23 Sep 2026Unrated
- 42.9%TTE-MatrixAgent + DeepSeek-V3.2 · 10 Nov 2025Emerging
Functions: Web researchProfessional deliverablesComputer useSpreadsheet workEnterprise workflowsFeature developmentLong-horizon autonomy
What we do not know
- Benchmarks measure tasks, not whole jobs. A job is a bundle of tasks, judgment, and relationships.
- Capability is not adoption. A task that agents can do in a benchmark may still be done by people at work.
- Most results come from single runs. Reliability over many runs is often unknown.
Occupations in this family
Sorted by US employment (BLS OEWS May 2025). O*NET sub-occupations share the employment figure of their SOC group.
How to cite
Credit the original work first: the benchmarks listed on this page, and their institutions.
Then, if you used this page:
Can Agents Work. "Can agents do office and administrative support work?." https://canagentswork.com/jobs/office-admin/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-jobs-office-admin,
title = {{Can agents do office and administrative support work?}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/jobs/office-admin/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}