Benchmarks / CRMArena-Pro
CRMArena-Pro
Built by Salesforce AI Research · Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal, et al. · released 24 May 2025
Salesforce AI Research's benchmark for agents that work inside a CRM. Agents answer sales, service, and configure-price-quote requests by querying two synthetic Salesforce orgs (one B2B, one B2C) through the Salesforce API, in single-turn and multi-turn settings with a simulated user.
Frontier
58.3%
Single-turn task success (B2C org)
gemini-2.5-pro (ReAct) · harness: ReAct
24 May 2025 · Source: Salesforce AI Research (benchmark maintainers)
Table 2 of the paper. B2B single-turn 54.1. Multi-turn 35.1 (B2B) and 30.0 (B2C). Workflow execution alone reached 83.0 (B2B) and 90.0 (B2C) single-turn. No newer primary results exist.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Task completion: whether the agent's final answer matches the gold answer for each query. The 19 task types cover four skills: workflow execution, policy compliance, text understanding, and database querying. A second track checks whether agents refuse to reveal confidential data. The headline is the single-turn success rate on the B2C org, averaged over the four skills. Multi-turn scores, where a persona-driven simulated user holds back details, are much lower.
Single-turn task success (B2C org): Share of single-turn queries on the B2C org where the agent's answer matches the gold answer, averaged over the four business skills. The paper also reports B2B and multi-turn results. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- api-tools, chat
- Grading
- automated-tests, llm-judge
- Tasks
- 4,280
- Contamination
- Queries, gold answers, and the org data are public on Hugging Face. The orgs hold synthetic data that gpt-4o generated from Salesforce schemas.
- Reuse
- CC BY-NC 4.0 (repository LICENSE.txt). Research use only. (cite-only)
Limits to keep in mind
- Published scores come from the May 2025 paper (o1, gpt-4o, Gemini 2.5, Llama 3.1 and 4). The Hugging Face leaderboard covers only the original CRMArena, so newer models are not tracked here. Source
- The org data is synthetic and generated by gpt-4o. 66.7% of CRM experts rated the B2B data as realistic and 62.3% rated the B2C data as realistic. Source
- Multi-turn runs use an LLM user simulator. A manual check of 20 trajectories found one simulator error (5%). gpt-4o also extracts answers and judges confidentiality refusals. Source
- Agents run through a ReAct scaffold with API access only. GUI access to the orgs was withdrawn after a Salesforce system update. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Single-turn task success (B2C org) | Date | Source |
|---|---|---|---|
| gpt-4o (ReAct) harness: ReAct | 29.2% | 24 May 2025 | Salesforce AI Research · primary |
| o1 (ReAct) harness: ReAct | 49.5% | 24 May 2025 | Salesforce AI Research · primary |
| gemini-2.5-pro (ReAct) · frontier harness: ReAct | 58.3% | 24 May 2025 | Salesforce AI Research · primary |
Timeline
- 24 May 2025 — Salesforce releases CRMArena-Pro. Source
Where it sits in the atlas
Go to the source
Last checked 23 Sep 2026 against 5 primary sources. See an error? Tell us.
How to cite
Credit the original work first: CRMArena-Pro by Salesforce AI Research (https://arxiv.org/abs/2505.18878).
Then, if you used this page:
Can Agents Work. "CRMArena-Pro: frontier results and sources." https://canagentswork.com/benchmarks/crmarena-pro/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-crmarena-pro,
title = {{CRMArena-Pro: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/crmarena-pro/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}