Benchmarks / Aider Polyglot leaderboard

Aider Polyglot leaderboard (Aider Polyglot)

Built by Aider · Paul Gauthier · released 21 Dec 2024

225 of the hardest Exercism practice exercises in C++, Go, Java, JavaScript, Python, and Rust, run through the aider coding tool. Paul Gauthier picked the exercises that a set of seven models in late 2024 mostly failed, so that top models would spread out between roughly 5% and 50%.

Frontier

88%

Percent correct

gpt-5 (high) · harness: aider

23 Aug 2025 · Source: Aider (benchmark maintainers)

Top row of the leaderboard on 2026-09-23. 198 of 225 correct, diff edit format, 91.6% correct edit format, $29.08 total cost, aider 0.86.2.dev. No newer entries since 2025-10-03.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Aider

Aider Polyglot leaderboard: Percent correct over time, 4 recorded results. 0%20%40%60%80%100%Jan 2025Apr 2025Jul 2025 o1-2024-12-17 (high): 61.7% (21 Dec 2024) Gemini 2.5 Pro Preview 05-06: 76.9% (7 May 2025) o3-pro (high): 84.9% (28 Jun 2025) gpt-5 (high): 88% (23 Aug 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/aider-polyglot/"><img src="https://canagentswork.com/og/benchmarks-aider-polyglot.png" width="600" height="315" alt="Aider Polyglot leaderboard: the best result is 88% (gpt-5 (high), 23 Aug 2025)." loading="lazy"></a>

Markdown:

[![Aider Polyglot leaderboard: the best result is 88% (gpt-5 (high), 23 Aug 2025).](https://canagentswork.com/og/benchmarks-aider-polyglot.png)](https://canagentswork.com/benchmarks/aider-polyglot/)

What it measures

Share of exercises where the model edits the starter code so that the hidden unit tests pass. The model gets the exercise text and the starter files. Runs allow two tries; the leaderboard reports the pass rate after each try and uses the second as the headline. It also reports how often the model used the requested edit format and the run cost.

Percent correct: Percentage of the 225 exercises completed correctly within two attempts. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
repo, cli
Grading
automated-tests
Tasks
225
Contamination
High risk. Exercism exercises and many solutions are public on GitHub and have been for years.
Reuse
Exercises are copyright Exercism and used under the Exercism tracks' open-source licenses. The benchmark repository states no license of its own. (cite-only)

Limits to keep in mind

  • Small, self-contained exercises with clear specs, not real codebases or issues. The benchmark measures code editing and instruction following inside one tool. Source
  • The leaderboard stopped growing in late 2025. The newest entry is dated 2025-10-03 and the page was last updated 2025-11-20, so models released since are absent. Source
  • Scores depend on aider's edit formats and prompts. A model that ignores the edit format loses points even if its code is right. Source
  • Designed to give headroom to 50%. With the top score at 88%, the benchmark has only 27 unsolved exercises left for the leader. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemPercent correctDateSource
gpt-5 (high) · frontier
harness: aider
88%23 Aug 2025Aider · primary
o3-pro (high)
harness: aider
84.9%28 Jun 2025Aider · primary
Gemini 2.5 Pro Preview 05-06
harness: aider
76.9%7 May 2025Aider · primary
o1-2024-12-17 (high)
harness: aider
61.7%21 Dec 2024Aider · primary

Timeline

  • 21 Dec 2024 — Aider launches the polyglot leaderboard. Source

Where it sits in the atlas

Software engineering

Feature development

Last checked 23 Sep 2026 against 4 primary sources. See an error? Tell us.

How to cite

Credit the original work first: Aider Polyglot leaderboard by Aider.

Then, if you used this page:

Can Agents Work. "Aider Polyglot leaderboard: frontier results and sources." https://canagentswork.com/benchmarks/aider-polyglot/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-aider-polyglot,
  title        = {{Aider Polyglot leaderboard: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/aider-polyglot/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}