Benchmarks / Aider Polyglot leaderboard
Aider Polyglot leaderboard (Aider Polyglot)
Built by Aider · Paul Gauthier · released 21 Dec 2024
225 of the hardest Exercism practice exercises in C++, Go, Java, JavaScript, Python, and Rust, run through the aider coding tool. Paul Gauthier picked the exercises that a set of seven models in late 2024 mostly failed, so that top models would spread out between roughly 5% and 50%.
Frontier
88%
Percent correct
gpt-5 (high) · harness: aider
23 Aug 2025 · Source: Aider (benchmark maintainers)
Top row of the leaderboard on 2026-09-23. 198 of 225 correct, diff edit format, 91.6% correct edit format, $29.08 total cost, aider 0.86.2.dev. No newer entries since 2025-10-03.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Share of exercises where the model edits the starter code so that the hidden unit tests pass. The model gets the exercise text and the starter files. Runs allow two tries; the leaderboard reports the pass rate after each try and uses the second as the headline. It also reports how often the model used the requested edit format and the run cost.
Percent correct: Percentage of the 225 exercises completed correctly within two attempts. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- repo, cli
- Grading
- automated-tests
- Tasks
- 225
- Contamination
- High risk. Exercism exercises and many solutions are public on GitHub and have been for years.
- Reuse
- Exercises are copyright Exercism and used under the Exercism tracks' open-source licenses. The benchmark repository states no license of its own. (cite-only)
Limits to keep in mind
- Small, self-contained exercises with clear specs, not real codebases or issues. The benchmark measures code editing and instruction following inside one tool. Source
- The leaderboard stopped growing in late 2025. The newest entry is dated 2025-10-03 and the page was last updated 2025-11-20, so models released since are absent. Source
- Scores depend on aider's edit formats and prompts. A model that ignores the edit format loses points even if its code is right. Source
- Designed to give headroom to 50%. With the top score at 88%, the benchmark has only 27 unsolved exercises left for the leader. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
- 21 Dec 2024 — Aider launches the polyglot leaderboard. Source
Where it sits in the atlas
Go to the source
- Website aider.chat
- Full leaderboard aider.chat
- Code github.com
- Announcement aider.chat
Last checked 23 Sep 2026 against 4 primary sources. See an error? Tell us.
How to cite
Credit the original work first: Aider Polyglot leaderboard by Aider.
Then, if you used this page:
Can Agents Work. "Aider Polyglot leaderboard: frontier results and sources." https://canagentswork.com/benchmarks/aider-polyglot/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-aider-polyglot,
title = {{Aider Polyglot leaderboard: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/aider-polyglot/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}