Benchmarks / SpreadsheetBench

SpreadsheetBench

Built by Renmin University of China · Zeyao Ma, Bohan Zhang, Jing Zhang, Jifan Yu, et al. · released 21 Jun 2024

912 spreadsheet manipulation questions taken from real Excel forum posts, each paired with the user's actual workbook. Workbooks have multiple tables, odd layouts, and non-text elements. A solution is checked like an online judge: it must work on several test-case spreadsheets with different values.

Frontier

83.1%

Overall pass@1 (V1, 912 questions)

Qingqiu Agent

23 Jun 2026 · Source: Renmin University of China (benchmark maintainers)

Verified entry at the top of the V1 Full (912) table in the leaderboard data file on 2026-09-23. JT AlphaData (CMCC JIUTIAN) is second at 77.85, dated 2026-08-15.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Renmin University of China

SpreadsheetBench: Overall pass@1 (V1, 912 questions) over time, 5 recorded results. 0%20%40%60%80%100%Jul 2025Oct 2025Jan 2026Apr 2026Jul 2026 ChatGPT Agent w/ .xlsx: 45.5% (17 Jul 2025) Shortcut.ai: 59.3% (16 Oct 2025) Gemini in Google Sheets: 70.5% (10 Mar 2026) WPS AI (Seed 2.0): 73.5% (16 Jun 2026) Qingqiu Agent: 83.1% (23 Jun 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/spreadsheetbench/"><img src="https://canagentswork.com/og/benchmarks-spreadsheetbench.png" width="600" height="315" alt="SpreadsheetBench: the best result is 83.1% (Qingqiu Agent, 23 Jun 2026)." loading="lazy"></a>

Markdown:

[![SpreadsheetBench: the best result is 83.1% (Qingqiu Agent, 23 Jun 2026).](https://canagentswork.com/og/benchmarks-spreadsheetbench.png)](https://canagentswork.com/benchmarks/spreadsheetbench/)

What it measures

Whether a system can produce the exact cell values a user asked for in a real spreadsheet, across cell-level and sheet-level edits. The headline is overall pass@1 on the full 912-question V1 set. The site also keeps a 400-question expert-verified V1 subset (released December 2025, best 99.25%) and SpreadsheetBench 2 (321 workflow tasks on financial modeling, debugging, and charts; best 50.86%).

Overall pass@1 (V1, 912 questions): Share of the 912 questions answered correctly across all 2,729 test-case spreadsheets, as listed on the official leaderboard. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
documents, cli
Grading
state-check
Tasks
912
Human reference
Four Excel experts on a 50-instruction subset (3 test cases each) scored 71.33% under the soft restriction and 62.00% under the hard restriction. GPT-4o scored 18.35% and 15.02% on the same measures in the paper.
Contamination
Questions were rewritten from forum posts by GPT-4 and annotators, and spreadsheet values were changed to build test cases, so exact forum solutions do not apply directly. The full data set and answers are public.
Reuse
CC BY-SA 4.0, stated in the paper's maintenance plan and the repo README. (open-other)

Limits to keep in mind

  • Some leaderboard rows are marked unverified. They come from outside evaluations by OpenAI and Microsoft, not from the maintainers' own runs. Source
  • The site's hard-coded "Top Score" box showed 70.48% on 2026-09-23 while its leaderboard data file listed 83.11%. See data/conflicts/spreadsheetbench-v1-top-score.yaml. Source
  • Most top entries are commercial spreadsheet products (WPS, Google Sheets, Univer) that do not disclose the model or scaffold. Source
  • The human baseline covers only 50 of 912 instructions, so it is not directly comparable with leaderboard scores. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemOverall pass@1 (V1, 912 questions)DateSource
Qingqiu Agent · frontier83.1%23 Jun 2026Renmin University of China · primary
WPS AI (Seed 2.0)73.5%16 Jun 2026Renmin University of China · primary
Gemini in Google Sheets70.5%10 Mar 2026Renmin University of China · primary
Shortcut.ai59.3%16 Oct 2025Renmin University of China · primary
ChatGPT Agent w/ .xlsx45.5%17 Jul 2025Renmin University of China · primary

Where sources disagree

The SpreadsheetBench site shows two different top scores for the 912-question V1 set. The static "Top Score (OVERALL)" box reads 70.48%. The leaderboard data file that the page loads lists Qingqiu Agent at 83.11% (verified, 2026-06-23). The box looks stale. We show the leaderboard value.

  • 83.11spreadsheetbench.github.io (primary) · Leaderboard data file, seen 2026-09-23. Entry is marked verified.
  • 70.48spreadsheetbench.github.io (primary) · Static statistics box on the V1 overview page, seen 2026-09-23. Matches the Gemini in Google Sheets row dated 2026-03-10.

We show 83.11. Status: open.

Timeline

  • 23 Jun 2026 — SpreadsheetBench V1 leader reaches 83.11%. Source
  • 21 Jun 2024 — SpreadsheetBench launches with 912 real spreadsheet questions. Source

Where it sits in the atlas

Office and administrative supportData and analytics (partial)Finance and accounting (partial)

Spreadsheet work

Last checked 23 Sep 2026 against 7 primary sources. See an error? Tell us.

How to cite

Credit the original work first: SpreadsheetBench by Renmin University of China (https://arxiv.org/abs/2406.14991).

Then, if you used this page:

Can Agents Work. "SpreadsheetBench: frontier results and sources." https://canagentswork.com/benchmarks/spreadsheetbench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-spreadsheetbench,
  title        = {{SpreadsheetBench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/spreadsheetbench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}