Models scored
26
evaluated
Modality
text
Category
communication
+2 more
Published
2025
arxiv.org
Citations
258
Semantic Scholar
Influential
40
citations
References
35
cited works
Venue
arXiv.org
published in
Abstract
Victor Barres, Honghua Dong, Soham Ray, Xujie Si, et al. (+1)
Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce $\tau^2$-bench, with four key contributions: 1) A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2) A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3) A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4) Fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, $\tau^2$-bench provides a controlled testbed for agents that must both reason effectively and guide user actions.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Claude Opus 4.6 | Anthropic | 92 |
| 02 | Claude Sonnet 4.6 | Anthropic | 92 |
| 03 | Claude Opus 4.5 | Anthropic | 89 |
| 04 | LongCat-Flash-Thinking-2601 | Meituan | 89 |
| 05 | Claude Haiku 4.5 | Anthropic | 83 |
| 06 | GPT-5.2 | OpenAI | 82 |
| 07 | GPT-5 | OpenAI | 81 |
| 08 | o3 | OpenAI | 80 |
| 09 | Nova 2 Omni | Amazon | 78 |
| 10 | GPT-5.1 | OpenAI | 78 |
| 11 | GPT-5.1 Instant | OpenAI | 78 |
| 12 | GPT-5.1 Thinking | OpenAI | 78 |
| 13 | Nova 2 Pro | Amazon | 78 |
| 14 | Nova 2 Lite | Amazon | 77 |
| 15 | LongCat-Flash-Lite | Meituan | 73 |
| 16 | Qwen3-235B-A22B-Thinking-2507 | Alibaba Cloud / Qwen Team | 72 |
| 17 | LongCat-Flash-Thinking | Meituan | 72 |
| 18 | Qwen3-235B-A22B-Instruct-2507 | Alibaba Cloud / Qwen Team | 71 |
| 19 | LongCat-Flash-Chat | Meituan | 71 |
| 20 | Kimi K2-Instruct-0905 | Moonshot AI | 71 |
| 21 | Kimi K2 Instruct | Moonshot AI | 71 |
| 22 | Qwen3-Next-80B-A3B-Thinking | Alibaba Cloud / Qwen Team | 68 |
| 23 | GPT-4o | OpenAI | 63 |
| 24 | Nemotron 3 Super (120B A12B) | NVIDIA | 63 |
| 25 | Qwen3-Next-80B-A3B-Instruct | Alibaba Cloud / Qwen Team | 57 |
| 26 | Nemotron 3 Nano (30B A3B) | NVIDIA | 57 |
26 of 26 models · score normalized 0–100 where available