Benchmarks / GDPval
GDPval
Built by OpenAI · Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, et al. · released 25 Sep 2025
Real work tasks from 44 occupations in the 9 US industries that add the most to GDP. Experienced professionals wrote each task from their own work. Blinded experts from the same occupation compare the AI deliverable with a deliverable made by a professional.
Frontier
84.9%
Wins plus ties against experts
GPT-5.5
Apr 2026 · Source: OpenAI (benchmark maintainers)
From the GPT-5.5 launch post, which does not say whether blinded experts or the automated grader produced the number. OpenAI is both the benchmark maintainer and the model developer. The same table gives GPT-5.4 83.0%, GPT-5.5 Pro 82.3%, and Gemini 3.1 Pro 67.3%.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Whether an AI deliverable is judged better than or as good as a professional's deliverable for the same request. Tasks come with reference files and ask for documents, slides, spreadsheets, diagrams, or media. The full set has 1,320 tasks (30 per occupation). A 220-task gold subset (5 per occupation) is public and is the set used for the headline numbers. Occupations include software developers, lawyers, accountants, nurses, financial managers, sales managers, editors, and mechanical engineers.
Wins plus ties against experts: Share of blinded pairwise comparisons in which experts rate the model deliverable better than (win) or as good as (tie) the professional's deliverable. Higher is better.
Status compares the frontier with human parity at 50%. At 50% wins plus ties the model matches the expert baseline on average.
Facts
- Grain
- Project level: whole projects judged by an acceptance standard
- Environment
- documents, cli
- Grading
- pairwise-human, human-expert
- Tasks
- 220
- Human reference
- A deliverable made by an industry professional (average 14 years of experience). Gold subset tasks take an expert 9.49 hours on average (median 5 hours) and are worth $398 on average at median wages.
- Contamination
- The gold subset is public, so later models may have seen it. OpenAI has not published a contamination analysis.
- Reuse
- The 220-task gold subset (prompts and reference files) is public on Hugging Face. The other 1,100 tasks and the expert grades are private. Rubrics and gold deliverables were later released. (cite-only)
Limits to keep in mind
- Tasks are one-shot and well specified. They do not test clarifying questions, iteration with a client, or building context over time. Source
- OpenAI builds the benchmark and reports its own models. Most later results come from OpenAI launch posts, which do not always say whether expert graders or the automated grader produced them. Source
- Human inter-rater agreement is 71%, so a single comparison is noisy. The automated grader agrees with humans 66% of the time and favors OpenAI outputs. Source
- Graders may have guessed which deliverable came from a model because of style cues, for example em dashes or first-person phrasing. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
Where it sits in the atlas
Management and business operationsFinance and accounting (partial)Legal (partial)Software engineering (partial)Sales and marketing (partial)Office and administrative support (partial)Healthcare (partial)Design, media, and writing (partial)Architecture and engineering (partial)Customer support (partial)Education and social services (partial)
Go to the source
- Website openai.com
- Paper arxiv.org
- Full leaderboard evals.openai.com
- Dataset huggingface.co
Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.
How to cite
Credit the original work first: GDPval by OpenAI (https://arxiv.org/abs/2510.04374).
Then, if you used this page:
Can Agents Work. "GDPval: frontier results and sources." https://canagentswork.com/benchmarks/gdpval/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-gdpval,
title = {{GDPval: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/gdpval/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}