Benchmarks / GDPval-AA

GDPval-AA

Built by Artificial Analysis · released Dec 2025

Artificial Analysis runs OpenAI's public 220-task GDPval gold set through its own agent harness, Stirrup, and ranks models by Elo from blind pairwise comparisons of the deliverables. The judges are a panel of three frontier language models, not human experts.

Frontier

1846 Elo

GDPval-AA Elo

Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) · harness: Stirrup

23 Sep 2026 · Source: Artificial Analysis (benchmark maintainers)

Rank 1 on the GDPval-AA v2.1 board, seen 2026-09-23. Confidence interval shown as -23 / +23. Model release month shown as Sep 2026. Judged by a panel of three frontier LLMs, not humans.

Unrated

No fixed reference point (for example Elo scores or field signals).

See the full leaderboard at Artificial Analysis

GDPval-AA: GDPval-AA Elo over time, 5 recorded results. 500 Elo1000 Elo1500 Elo2000 Elo2500 EloSep 2026Sep 2026Sep 2026Sep 2026Oct 2026Oct 2026 GPT-5 (high): 906 Elo (23 Sep 2026) Claude Opus 4.7 (Adaptive Reasoning, Max Effort): 1338 Elo (23 Sep 2026) Claude Opus 5 (Adaptive Reasoning, Max Effort): 1708 Elo (23 Sep 2026) GPT-5.6 Sol (max): 1588 Elo (23 Sep 2026) Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback): 1846 Elo (23 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/gdpval-aa/"><img src="https://canagentswork.com/og/benchmarks-gdpval-aa.png" width="600" height="315" alt="GDPval-AA: the best result is 1846 Elo (Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback), 23 Sep 2026)." loading="lazy"></a>

Markdown:

[![GDPval-AA: the best result is 1846 Elo (Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback), 23 Sep 2026).](https://canagentswork.com/og/benchmarks-gdpval-aa.png)](https://canagentswork.com/benchmarks/gdpval-aa/)

What it measures

How well a model, acting as an agent with a Linux sandbox, code execution, web search, web fetch, and image viewing, produces work deliverables (documents, slides, spreadsheets, media) for the 220 GDPval gold tasks across 44 occupations. A judge sampled from a panel of three frontier LLMs blindly picks the better of two submissions to the same task. Ratings are fitted with a Crowd-BT model and reported as Elo with 95% confidence intervals.

GDPval-AA Elo: Crowd-BT rating from blind pairwise LLM-judge comparisons of deliverables. Anchored to DeepSeek V4.1 Flash (max) at 1600 in v2.1. Ties count as half a win for each side. Higher is better.

No fixed reference point, so we do not rate its status. Elo scale with a model anchor. Scores are relative and shift when the anchor changes.

Facts

Grain
Project level: whole projects judged by an acceptance standard
Environment
cli, api-tools, documents
Grading
llm-judge
Tasks
220
Human reference
None on the current v2.1 scale. In GDPval-AA v2 (June to September 2026) the scale was anchored so that human expert performance was 1000 Elo. v2.1 re-anchored the scale to DeepSeek V4.1 Flash (max) at 1600, so the human anchor no longer applies.
Contamination
Uses the public GDPval gold subset, so models may have seen the tasks. Artificial Analysis has not published a contamination analysis.
Reuse
Tasks are OpenAI's public GDPval gold set. The Stirrup harness is open source on GitHub. Artificial Analysis publishes Elo scores and example submissions on its site. (cite-only)

Limits to keep in mind

  • Judges are language models (Claude Opus 5, GPT-5.6 Sol, Gemini 3.8 Flash in September 2026), not the human experts used in OpenAI's own GDPval grading. Source
  • The Elo scale has been re-anchored twice (v2 in June 2026 to human experts at 1000, v2.1 in September 2026 to DeepSeek V4.1 Flash at 1600). Scores from different versions are not comparable. Source
  • Each model runs once per task, and 95% confidence intervals are about plus or minus 20 to 30 Elo, so nearby ranks can swap. Source
  • Many entries are the same model at different reasoning-effort settings, so the board mixes model quality with compute budget. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemGDPval-AA EloDateSource
GPT-5 (high)
harness: Stirrup
906 Elo23 Sep 2026Artificial Analysis · primary
Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
harness: Stirrup
1338 Elo23 Sep 2026Artificial Analysis · primary
Claude Opus 5 (Adaptive Reasoning, Max Effort)
harness: Stirrup
1708 Elo23 Sep 2026Artificial Analysis · primary
GPT-5.6 Sol (max)
harness: Stirrup
1588 Elo23 Sep 2026Artificial Analysis · primary
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) · frontier
harness: Stirrup
1846 Elo23 Sep 2026Artificial Analysis · primary

Timeline

  • Sep 2026 — GDPval-AA v2.1 re-anchors the Elo scale. Source
  • Dec 2025 — Artificial Analysis launches GDPval-AA. Source

Where it sits in the atlas

Management and business operationsFinance and accounting (partial)Legal (partial)Software engineering (partial)Sales and marketing (partial)Office and administrative support (partial)Healthcare (partial)Design, media, and writing (partial)Architecture and engineering (partial)

Professional deliverables

Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.

How to cite

Credit the original work first: GDPval-AA by Artificial Analysis (https://arxiv.org/abs/2510.04374).

Then, if you used this page:

Can Agents Work. "GDPval-AA: frontier results and sources." https://canagentswork.com/benchmarks/gdpval-aa/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-gdpval-aa,
  title        = {{GDPval-AA: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/gdpval-aa/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}