Benchmarks / Finance Agent Benchmark

Finance Agent Benchmark (Finance Agent)

Built by Vals AI · Antoine Bigeard, Langston Nashold, Rayan Krishnan, Shirley Wu · released 20 May 2025

Vals AI's test of whether an agent can do the research work of an entry-level financial analyst. Each question asks about public companies and their SEC filings, and the agent must find the answer with EDGAR search, web search, a page parser, and a retrieval tool.

Frontier

64.4%

Accuracy

Claude Opus 4.7 · harness: Vals finance-agent v1.1

4 Jun 2026 · Source: Vals AI (benchmark maintainers)

Top of the v1.1 leaderboard (page data 64.373, stderr 2.79, cost about $0.80 per question). Muse Spark 60.59, DeepSeek V4 Pro 60.39, and Claude Opus 4.6 (Thinking) 60.05 follow. Date is the page's "updated" field, not the run date.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Vals AI

Finance Agent Benchmark: Accuracy over time, 4 recorded results. 0%20%40%60%80%100%Jul 2025Oct 2025Jan 2026Apr 2026 o3: 46.8% (20 May 2025) GPT-5.2: 58.5% (4 Jun 2026) Claude Sonnet 4.6: 63.3% (4 Jun 2026) Claude Opus 4.7: 64.4% (4 Jun 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/vals-finance-agent/"><img src="https://canagentswork.com/og/benchmarks-vals-finance-agent.png" width="600" height="315" alt="Finance Agent Benchmark: the best result is 64.4% (Claude Opus 4.7, 4 Jun 2026)." loading="lazy"></a>

Markdown:

[![Finance Agent Benchmark: the best result is 64.4% (Claude Opus 4.7, 4 Jun 2026).](https://canagentswork.com/og/benchmarks-vals-finance-agent.png)](https://canagentswork.com/benchmarks/vals-finance-agent/)

What it measures

Final-answer accuracy on a private test set of 337 expert-written questions (537 in total, with 50 public and 150 licensable validation questions). Questions span nine categories from simple retrieval to financial modeling and market analysis. An LLM judge (GPT-5.2, mode of three runs) compares each answer with the expert answer. Vals also records cost, latency, and tool calls.

Accuracy: Share of private test questions where the LLM judge accepts the agent's final answer. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
api-tools, documents
Grading
llm-judge
Tasks
537
Human reference
Finance experts from banks, private equity firms, and hedge funds wrote and answered the questions. No human accuracy score is published.
Contamination
The test set is private and Vals says it will stay private.
Reuse
50 public validation questions in the repository (MIT). 150 validation questions are available for license. The 337-question test set is private. (cite-only)

Limits to keep in mind

  • Version 1.1 (early 2026) changed the data, tools, prompts, and judge, and re-ran every model. Scores from version 1.0, including the paper's 46.8% for o3, are not comparable. Source
  • Each score has a standard error of about 2.8 points, so the top five models (60% to 64%) overlap. Source
  • Grading uses an LLM judge. Vals takes the mode of three GPT-5.2 judgments to reduce variance. Source
  • The agent only sees the fixed tool set in the Vals harness. Access to the Vals platform to run the private set is gated. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemAccuracyDateSource
GPT-5.2
harness: Vals finance-agent v1.1
58.5%4 Jun 2026Vals AI · primary
Claude Sonnet 4.6
harness: Vals finance-agent v1.1
63.3%4 Jun 2026Vals AI · primary
Claude Opus 4.7 · frontier
harness: Vals finance-agent v1.1
64.4%4 Jun 2026Vals AI · primary
o3
harness: Vals finance-agent v1.0
46.8%20 May 2025Vals AI · primary

Timeline

  • 20 May 2025 — Vals AI releases the Finance Agent Benchmark. Source

Where it sits in the atlas

Finance and accounting

Financial analysisWeb research

Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.

How to cite

Credit the original work first: Finance Agent Benchmark by Vals AI (https://arxiv.org/abs/2508.00828).

Then, if you used this page:

Can Agents Work. "Finance Agent Benchmark: frontier results and sources." https://canagentswork.com/benchmarks/vals-finance-agent/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-vals-finance-agent,
  title        = {{Finance Agent Benchmark: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/vals-finance-agent/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}