Benchmarks / GDPval

GDPval

Built by OpenAI · Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, et al. · released 25 Sep 2025

Real work tasks from 44 occupations in the 9 US industries that add the most to GDP. Experienced professionals wrote each task from their own work. Blinded experts from the same occupation compare the AI deliverable with a deliverable made by a professional.

Frontier

84.9%

Wins plus ties against experts

GPT-5.5

Apr 2026 · Source: OpenAI (benchmark maintainers)

From the GPT-5.5 launch post, which does not say whether blinded experts or the automated grader produced the number. OpenAI is both the benchmark maintainer and the model developer. The same table gives GPT-5.4 83.0%, GPT-5.5 Pro 82.3%, and Gemini 3.1 Pro 67.3%.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at OpenAI

GDPval: Wins plus ties against experts over time, 4 recorded results. 0%20%40%60%80%100%Oct 2025Jan 2026Apr 2026human parity Claude Opus 4.1: 47.6% (Sep 2025) GPT-5.2 Thinking: 70.9% (Dec 2025) Claude Opus 4.7: 80.3% (Apr 2026) GPT-5.5: 84.9% (Apr 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/gdpval/"><img src="https://canagentswork.com/og/benchmarks-gdpval.png" width="600" height="315" alt="GDPval: the best result is 84.9% (GPT-5.5, Apr 2026)." loading="lazy"></a>

Markdown:

[![GDPval: the best result is 84.9% (GPT-5.5, Apr 2026).](https://canagentswork.com/og/benchmarks-gdpval.png)](https://canagentswork.com/benchmarks/gdpval/)

What it measures

Whether an AI deliverable is judged better than or as good as a professional's deliverable for the same request. Tasks come with reference files and ask for documents, slides, spreadsheets, diagrams, or media. The full set has 1,320 tasks (30 per occupation). A 220-task gold subset (5 per occupation) is public and is the set used for the headline numbers. Occupations include software developers, lawyers, accountants, nurses, financial managers, sales managers, editors, and mechanical engineers.

Wins plus ties against experts: Share of blinded pairwise comparisons in which experts rate the model deliverable better than (win) or as good as (tie) the professional's deliverable. Higher is better.

Status compares the frontier with human parity at 50%. At 50% wins plus ties the model matches the expert baseline on average.

Facts

Grain
Project level: whole projects judged by an acceptance standard
Environment
documents, cli
Grading
pairwise-human, human-expert
Tasks
220
Human reference
A deliverable made by an industry professional (average 14 years of experience). Gold subset tasks take an expert 9.49 hours on average (median 5 hours) and are worth $398 on average at median wages.
Contamination
The gold subset is public, so later models may have seen it. OpenAI has not published a contamination analysis.
Reuse
The 220-task gold subset (prompts and reference files) is public on Hugging Face. The other 1,100 tasks and the expert grades are private. Rubrics and gold deliverables were later released. (cite-only)

Limits to keep in mind

  • Tasks are one-shot and well specified. They do not test clarifying questions, iteration with a client, or building context over time. Source
  • OpenAI builds the benchmark and reports its own models. Most later results come from OpenAI launch posts, which do not always say whether expert graders or the automated grader produced them. Source
  • Human inter-rater agreement is 71%, so a single comparison is noisy. The automated grader agrees with humans 66% of the time and favors OpenAI outputs. Source
  • Graders may have guessed which deliverable came from a model because of style cues, for example em dashes or first-person phrasing. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemWins plus ties against expertsDateSource
Claude Opus 4.780.3%Apr 2026OpenAI · primary
GPT-5.5 · frontier84.9%Apr 2026OpenAI · primary
GPT-5.2 Thinking70.9%Dec 2025OpenAI · primary
Claude Opus 4.147.6%Sep 2025OpenAI · primary

Timeline

  • Apr 2026 — GPT-5.5 reaches 84.9% wins plus ties on GDPval. Source
  • Dec 2025 — GPT-5.2 Thinking passes the 50% line on GDPval. Source
  • 25 Sep 2025 — OpenAI releases GDPval. Source

Where it sits in the atlas

Management and business operationsFinance and accounting (partial)Legal (partial)Software engineering (partial)Sales and marketing (partial)Office and administrative support (partial)Healthcare (partial)Design, media, and writing (partial)Architecture and engineering (partial)Customer support (partial)Education and social services (partial)

Professional deliverables

Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.

How to cite

Credit the original work first: GDPval by OpenAI (https://arxiv.org/abs/2510.04374).

Then, if you used this page:

Can Agents Work. "GDPval: frontier results and sources." https://canagentswork.com/benchmarks/gdpval/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-gdpval,
  title        = {{GDPval: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/gdpval/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}