Benchmarks / SpreadsheetBench 2
SpreadsheetBench 2
Built by Renmin University of China, Aptura, AfterQuery and Shortcut · Jian Zhu, Yuzheng Zhang, Zeyao Ma, Bohan Zhang, et al. · released 29 Jun 2026
321 business spreadsheet tasks built from real financial reports and corporate filings and checked by domain experts: complete a template, build out a financial model, find and fix errors in a workbook, or draw charts. Workbooks average 11.8 sheets, and a task needs 593.5 cell changes on average. The successor to SpreadsheetBench from the same Renmin University group.
Frontier
50.9%
Overall accuracy
WPS AI
29 Aug 2026 · Source: Renmin University of China (benchmark maintainers)
Rank 1 on the V2 full leaderboard. Commercial product by Kingsoft Office; model and scaffold not disclosed. Marked verified by the maintainers. Subscores: template 76.29, financial model 56.0, debugging 16.0, visualization 71.94. The site's static top-score box still shows 34.89%.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Share of tasks solved. For template, financial-model, and debugging tasks the output workbook must match the golden file in every target cell with no other cell changed. For visualization tasks a vision-language model checks the chart against a checklist of expert assertions and reports the share passed. The overall score aggregates the four categories (template, financial model, debugging, visualization). Rows marked verified were run or checked by the maintainers; scaffolded models use a SWE-agent-based scaffold.
Overall accuracy: Share of the 321 tasks solved, as listed in the site's V2 full leaderboard, with subscores for template, financial model, debugging, and visualization. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- documents, cli
- Grading
- state-check, llm-judge
- Tasks
- 321
- Contamination
- Tasks and golden files are public on Hugging Face. The maintainers have not published a contamination analysis. Results are submitted by email with logs and outputs and checked by the maintainers, who process submissions weekly.
- Reuse
- MIT on the Hugging Face dataset card (KAKA22/SpreadsheetBench-v2). The code repository has no license file. (open-mit)
Limits to keep in mind
- The site's static "Top Score (OVERALL)" box for V2 showed 34.89% on 2026-09-24 while the leaderboard data file listed WPS AI at 50.86%. See data/conflicts/spreadsheetbench-2-top-score.yaml. Source
- The two leading rows are commercial spreadsheet products (WPS AI, arito) that do not disclose the model or scaffold. The best scaffolded model is Claude Opus 4.6 with SWE-agent at 34.89%. Source
- Visualization tasks are scored by a vision-language model against assertion checklists, and the site gives two readings of that subscore (average assertion pass rate, or a task counted correct above 70), so the overall score mixes exact-match completion with rubric credit. Source
- Debugging is the weak spot: the best debugging subscore is 28.0% (arito) and the leader WPS AI scores 16.0% there. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Overall accuracy | Date | Source |
|---|---|---|---|
| WPS AI · frontier | 50.9% | 29 Aug 2026 | Renmin University of China · primary |
| arito | 45.5% | 4 Aug 2026 | Renmin University of China · primary |
| Claude Opus 4.6 (SWE-agent) harness: SWE-agent | 34.9% | 29 Jun 2026 | Renmin University of China · primary |
Where sources disagree
The SpreadsheetBench site shows two different top scores for V2. The static "Top Score (OVERALL)" box on the V2 overview reads 34.89%, the best scaffolded model in the paper. The leaderboard data file that the page loads lists WPS AI at 50.86% (verified, 2026-08-29). The box looks stale. We show the leaderboard value.
- 50.86 — spreadsheetbench.github.io (primary) · Leaderboard data file, seen 2026-09-24. Entry is marked verified.
- 34.89 — spreadsheetbench.github.io (primary) · Static statistics box on the V2 overview page, seen 2026-09-24. Matches the Claude Opus 4.6 (SWE-agent) row and the paper abstract.
We show 50.86. Status: open.
Timeline
Where it sits in the atlas
Work ladder: Direct evidence for Finance and accounting. On the work ladder it counts as 50.9% of tasks completed. How the ladder works
Work it measures (O*NET work activities): Analyze business or financial data; Develop financial or business plans; Create visual designs or displays.
Finance and accountingOffice and administrative support (partial)Data and analytics (partial)
Go to the source
- Website spreadsheetbench.github.io
- Paper arxiv.org
- Full leaderboard spreadsheetbench.github.io
- Code github.com
- Dataset huggingface.co
Last checked 24 Sep 2026 against 5 primary sources. See an error? Tell us.
How to cite
Credit the original work first: SpreadsheetBench 2 by Renmin University of China, Aptura, AfterQuery, and Shortcut (https://arxiv.org/abs/2606.29955).
Then, if you used this page:
Can Agents Work. "SpreadsheetBench 2: frontier results and sources." https://canagentswork.com/benchmarks/spreadsheetbench-2/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-spreadsheetbench-2,
title = {{SpreadsheetBench 2: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/spreadsheetbench-2/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}