Benchmarks / τ²-bench
τ²-bench
Built by Sierra · Victor Barres, Honghua Dong, Soham Ray, Xujie Si, et al. · released 9 Jun 2025
Simulated customer-service conversations from Sierra. An agent chats with an LLM-played user, calls tools, and must follow a written domain policy. The telecom domain adds dual control: the user also has tools, so the agent must guide the user through steps that only the user can perform.
Frontier
97.8%
Telecom pass^1
Qwen3.5-397B-A17B (thinking) · harness: tau2-bench default agent
27 Feb 2026 · Source: Sierra (benchmark maintainers)
Best telecom pass^1 among runs that Sierra ran and verified with trajectories. User simulator gpt-5.2 (low reasoning), 4 trials, self-hosted with vLLM. pass^2 95.8, pass^3 93.9, pass^4 92.1. Same submission: airline 81.5, retail 84.4, banking knowledge 9.8 pass^1.
The best result is at 90% or more of the ceiling.
What it measures
Whether the final database state and required outputs match the expected result for each task. Each task runs several times. pass^k is the chance that all k tries of a task succeed, so higher k rewards reliability. Domains are airline (50 tasks), retail (114), and telecom (114). We use telecom pass^1 as the headline because telecom is the domain that τ²-bench introduced. The repository now ships as τ³-bench with a banking knowledge domain and a voice mode; those are not scored here.
Telecom pass^1: Share of telecom tasks that the agent completes correctly on a single try, averaged over trials. The leaderboard reports pass^1 to pass^4 per domain. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- api-tools, chat
- Grading
- state-check
- Tasks
- 278
- Contamination
- Tasks, policies, and tools are public in the repository. The leaderboard labels models trained on τ-bench domains or tasks as custom submissions, but it cannot detect training on the public data.
- Reuse
- MIT (repository license covers tasks, policies, and code). (open-mit)
Limits to keep in mind
- The user is an LLM simulator. The choice of user model changes scores, and submissions use different user models (gpt-4.1, gpt-5.2, or the agent's own model). Sierra recommends gpt-5.2. Source
- Telecom pass^1 is close to the ceiling: several models score 97-98%. pass^4 and the newer banking knowledge domain (best 55.2% pass^1) still separate models. Source
- In February 2026 Sierra fixed 50+ airline and retail tasks. Airline pass^1 rose by 14 to 20 points after the fixes, so airline and retail scores from before and after are not comparable. Source
- Some leaderboard entries were submitted by model developers with modified prompts and no trajectories. Sierra marks these as unverified in the submission files. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Telecom pass^1 | Date | Source |
|---|---|---|---|
| Qwen3.5-397B-A17B (thinking) · frontier harness: tau2-bench default agent | 97.8% | 27 Feb 2026 | Sierra · primary |
| Qwen3-Max-Thinking harness: tau2-bench default agent | 98.2% | 23 Jan 2026 | Qwen team (leaderboard submission) · lab-reported |
| GPT-5 harness: tau2-bench default agent | 95.8% | 9 Aug 2025 | Sierra · primary |
| GPT-4.1 harness: tau2-bench default agent | 34% | 9 Jun 2025 | Sierra · primary |
| o4-mini harness: tau2-bench default agent | 50.2% | 9 Jun 2025 | Sierra · primary |
Timeline
Where it sits in the atlas
Go to the source
- Website taubench.com
- Paper arxiv.org
- Full leaderboard taubench.com
- Code github.com
- Announcement sierra.ai
Last checked 23 Sep 2026 against 8 primary sources. See an error? Tell us.
How to cite
Credit the original work first: τ²-bench by Sierra (https://arxiv.org/abs/2506.07982).
Then, if you used this page:
Can Agents Work. "τ²-bench: frontier results and sources." https://canagentswork.com/benchmarks/tau2-bench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-tau2-bench,
title = {{τ²-bench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/tau2-bench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}