Benchmarks / GDPval-AA
GDPval-AA
Built by Artificial Analysis · released Dec 2025
Artificial Analysis runs OpenAI's public 220-task GDPval gold set through its own agent harness, Stirrup, and ranks models by Elo from blind pairwise comparisons of the deliverables. The judges are a panel of three frontier language models, not human experts.
Frontier
1846 Elo
GDPval-AA Elo
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) · harness: Stirrup
23 Sep 2026 · Source: Artificial Analysis (benchmark maintainers)
Rank 1 on the GDPval-AA v2.1 board, seen 2026-09-23. Confidence interval shown as -23 / +23. Model release month shown as Sep 2026. Judged by a panel of three frontier LLMs, not humans.
No fixed reference point (for example Elo scores or field signals).
What it measures
How well a model, acting as an agent with a Linux sandbox, code execution, web search, web fetch, and image viewing, produces work deliverables (documents, slides, spreadsheets, media) for the 220 GDPval gold tasks across 44 occupations. A judge sampled from a panel of three frontier LLMs blindly picks the better of two submissions to the same task. Ratings are fitted with a Crowd-BT model and reported as Elo with 95% confidence intervals.
GDPval-AA Elo: Crowd-BT rating from blind pairwise LLM-judge comparisons of deliverables. Anchored to DeepSeek V4.1 Flash (max) at 1600 in v2.1. Ties count as half a win for each side. Higher is better.
No fixed reference point, so we do not rate its status. Elo scale with a model anchor. Scores are relative and shift when the anchor changes.
Facts
- Grain
- Project level: whole projects judged by an acceptance standard
- Environment
- cli, api-tools, documents
- Grading
- llm-judge
- Tasks
- 220
- Human reference
- None on the current v2.1 scale. In GDPval-AA v2 (June to September 2026) the scale was anchored so that human expert performance was 1000 Elo. v2.1 re-anchored the scale to DeepSeek V4.1 Flash (max) at 1600, so the human anchor no longer applies.
- Contamination
- Uses the public GDPval gold subset, so models may have seen the tasks. Artificial Analysis has not published a contamination analysis.
- Reuse
- Tasks are OpenAI's public GDPval gold set. The Stirrup harness is open source on GitHub. Artificial Analysis publishes Elo scores and example submissions on its site. (cite-only)
Limits to keep in mind
- Judges are language models (Claude Opus 5, GPT-5.6 Sol, Gemini 3.8 Flash in September 2026), not the human experts used in OpenAI's own GDPval grading. Source
- The Elo scale has been re-anchored twice (v2 in June 2026 to human experts at 1000, v2.1 in September 2026 to DeepSeek V4.1 Flash at 1600). Scores from different versions are not comparable. Source
- Each model runs once per task, and 95% confidence intervals are about plus or minus 20 to 30 Elo, so nearby ranks can swap. Source
- Many entries are the same model at different reasoning-effort settings, so the board mixes model quality with compute budget. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | GDPval-AA Elo | Date | Source |
|---|---|---|---|
| GPT-5 (high) harness: Stirrup | 906 Elo | 23 Sep 2026 | Artificial Analysis · primary |
| Claude Opus 4.7 (Adaptive Reasoning, Max Effort) harness: Stirrup | 1338 Elo | 23 Sep 2026 | Artificial Analysis · primary |
| Claude Opus 5 (Adaptive Reasoning, Max Effort) harness: Stirrup | 1708 Elo | 23 Sep 2026 | Artificial Analysis · primary |
| GPT-5.6 Sol (max) harness: Stirrup | 1588 Elo | 23 Sep 2026 | Artificial Analysis · primary |
| Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) · frontier harness: Stirrup | 1846 Elo | 23 Sep 2026 | Artificial Analysis · primary |
Timeline
Where it sits in the atlas
Management and business operationsFinance and accounting (partial)Legal (partial)Software engineering (partial)Sales and marketing (partial)Office and administrative support (partial)Healthcare (partial)Design, media, and writing (partial)Architecture and engineering (partial)
Go to the source
- Website artificialanalysis.ai
- Paper arxiv.org
- Full leaderboard artificialanalysis.ai
- Code github.com
- Announcement artificialanalysis.ai
- Dataset huggingface.co
Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.
How to cite
Credit the original work first: GDPval-AA by Artificial Analysis (https://arxiv.org/abs/2510.04374).
Then, if you used this page:
Can Agents Work. "GDPval-AA: frontier results and sources." https://canagentswork.com/benchmarks/gdpval-aa/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-gdpval-aa,
title = {{GDPval-AA: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/gdpval-aa/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}