Models scored
35
evaluated
Modality
text
Category
communication
+2 more
Published
2025
arxiv.org
Citations
258
Semantic Scholar
Influential
40
citations
References
35
cited works
Venue
arXiv.org
published in
Abstract
Victor Barres, Honghua Dong, Soham Ray, Xujie Si, et al. (+1)
Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce $\tau^2$-bench, with four key contributions: 1) A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2) A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3) A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4) Fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, $\tau^2$-bench provides a controlled testbed for agents that must both reason effectively and guide user actions.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | LongCat-Flash-Thinking-2601 | Meituan | 99 |
| 02 | Claude Opus 4.6 | Anthropic | 99 |
| 03 | GPT-5.4 | OpenAI | 99 |
| 04 | GPT-5.2 | OpenAI | 99 |
| 05 | Claude Opus 4.5 | Anthropic | 98 |
| 06 | GPT-5.5 | OpenAI | 98 |
| 07 | Claude Sonnet 4.6 | Anthropic | 98 |
| 08 | MiMo-V2-Pro | Xiaomi | 97 |
| 09 | GPT-5 | OpenAI | 97 |
| 10 | GPT-5.1 Thinking | OpenAI | 96 |
| 11 | GPT-5.1 | OpenAI | 96 |
| 12 | GPT-5.1 Instant | OpenAI | 96 |
| 13 | GPT-5.4 mini | OpenAI | 93 |
| 14 | Nova 2 Pro | Amazon | 93 |
| 15 | GPT-5.4 nano | OpenAI | 93 |
| 16 | Muse Spark | Meta | 92 |
| 17 | MiniMax M2 | MiniMax | 87 |
| 18 | MiniMax M2.1 | MiniMax | 87 |
| 19 | Command A+ | Cohere | 85 |
| 20 | LongCat-Flash-Thinking | Meituan | 83 |
| 21 | Claude Haiku 4.5 | Anthropic | 83 |
| 22 | Nova 2 Omni | Amazon | 80 |
| 23 | Nova 2 Lite | Amazon | 76 |
| 24 | LongCat-Flash-Chat | Meituan | 74 |
| 25 | LongCat-Flash-Lite | Meituan | 73 |
| 26 | MAI-Code-1-Flash | Microsoft | 72 |
| 27 | Kimi K2 Instruct | Moonshot AI | 66 |
| 28 | Kimi K2-Instruct-0905 | Moonshot AI | 66 |
| 29 | Nemotron 3 Super (120B A12B) | NVIDIA | 64 |
| 30 | o3 | OpenAI | 58 |
| 31 | Qwen3-235B-A22B-Thinking-2507 | Alibaba Cloud / Qwen Team | 46 |
| 32 | Qwen3-Next-80B-A3B-Thinking | Alibaba Cloud / Qwen Team | 44 |
| 33 | Nemotron 3 Nano (30B A3B) | NVIDIA | 42 |
| 34 | GPT-4o | OpenAI | 24 |
| 35 | Qwen3-Next-80B-A3B-Instruct | Alibaba Cloud / Qwen Team | 13 |
35 of 35 models · score normalized 0–100 where available