Benchmarks / Vending-Bench 2
Vending-Bench 2
Built by Andon Labs · released 18 Nov 2025
An agent runs a simulated vending machine business for one simulated year, starting with $500. It finds suppliers on a simulated web, negotiates by email, orders and stocks products, sets prices, and handles delays, bankrupt suppliers, refund demands, and suppliers that try to cheat it.
Frontier
$15,515
Final bank balance
GPT-6 Astra
7 Sep 2026 · Source: Andon Labs (benchmark maintainers)
Leaderboard value (± $1,074, "average across 5 runs"). The 2026-09-07 blog post gives $15,515 as the average of six runs; the blog also says every Astra run beat every Claude Fable 5.1 run (Fable 5.1 average $5,422).
No fixed reference point (for example Elo scores or field signals).
What it measures
The bank balance at the end of the year, averaged over several runs. The agent pays a $2 daily fee and is terminated early if it cannot pay for more than 10 days in a row. Sales depend on day of week, season, weather, and price. A full run produces 3,000 to 6,000 messages and 60 to 100 million output tokens, so the score mostly reflects whether the agent stays coherent and keeps negotiating well over a long horizon.
Final bank balance: Money balance after one simulated year, averaged across runs. There is no fixed ceiling. Higher is better.
No fixed reference point, so we do not rate its status. No ceiling by design. Andon Labs' own rough estimate of a good human-level strategy is about $63k, roughly four times the best model score in September 2026.
Facts
- Grain
- Project level: whole projects judged by an acceptance standard
- Environment
- simulated-business
- Grading
- outcome-metric
- Human reference
- No human baseline was run. Andon Labs estimates that a "good" strategy (best-selling product, half-price supply, optimal stocking) would make about $206 per day, or roughly $63k in a year.
- Contamination
- The environment is not public, but the system prompt is shown on the benchmark page and the sales simulation follows the equations in the original Vending-Bench paper.
- Reuse
- No public code or data release is linked from the benchmark page. Andon Labs runs the simulation. (cite-only)
Limits to keep in mind
- Run-to-run variance is large. Leaderboard entries carry error bars of roughly $370 to $2,100 on averages of five runs. Source
- Suppliers are other LLMs. The page notes they could in theory be jailbroken into giving stock away, and the sales equations can be gamed. Source
- Each run uses 60 to 100 million output tokens, so results are costly to reproduce and only Andon Labs reports them. Source
- The score ignores conduct. Andon Labs' blog posts describe top models that lied to customers, fixed prices with competitors, or withheld refunds while still scoring well. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Final bank balance | Date | Source |
|---|---|---|---|
| GPT-6 Astra · frontier | $15,515 | 7 Sep 2026 | Andon Labs · primary |
| Claude Opus 5 | $11,182 | 28 Jul 2026 | Andon Labs · primary |
| Claude Opus 4.6 | $8,018 | 4 Feb 2026 | Andon Labs · primary |
| Gemini 3 Pro | $5,478 | 18 Nov 2025 | Andon Labs · primary |
Timeline
Where it sits in the atlas
Go to the source
- Website andonlabs.com
- Paper arxiv.org
- Full leaderboard andonlabs.com
- Announcement andonlabs.com
Last checked 23 Sep 2026 against 5 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: Vending-Bench 2 by Andon Labs (https://arxiv.org/abs/2502.15840).
Then, if you used this page:
Can Agents Work. "Vending-Bench 2: frontier results and sources." https://canagentswork.com/benchmarks/vending-bench-2/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-vending-bench-2,
title = {{Vending-Bench 2: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/vending-bench-2/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}