Tau2 Telecom

Evaluating Conversational Agents in a Dual-Control Environment

Models scored

35

evaluated

Modality

text

Category

communication

+2 more

Published

2025

arxiv.org

Citations

258

Semantic Scholar

Influential

40

citations

References

35

cited works

Venue

arXiv.org

published in

Abstract

Victor Barres, Honghua Dong, Soham Ray, Xujie Si, et al. (+1)

Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce $\tau^2$-bench, with four key contributions: 1) A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2) A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3) A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4) Fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, $\tau^2$-bench provides a controlled testbed for agents that must both reason effectively and guide user actions.

communicationreasoningtool calling

Search

#ModelLabScore
01LongCat-Flash-Thinking-2601Meituan99
02Claude Opus 4.6Anthropic99
03GPT-5.4OpenAI99
04GPT-5.2OpenAI99
05Claude Opus 4.5Anthropic98
06GPT-5.5OpenAI98
07Claude Sonnet 4.6Anthropic98
08MiMo-V2-ProXiaomi97
09GPT-5OpenAI97
10GPT-5.1 ThinkingOpenAI96
11GPT-5.1OpenAI96
12GPT-5.1 InstantOpenAI96
13GPT-5.4 miniOpenAI93
14Nova 2 ProAmazon93
15GPT-5.4 nanoOpenAI93
16Muse SparkMeta92
17MiniMax M2MiniMax87
18MiniMax M2.1MiniMax87
19Command A+Cohere85
20LongCat-Flash-ThinkingMeituan83
21Claude Haiku 4.5Anthropic83
22Nova 2 OmniAmazon80
23Nova 2 LiteAmazon76
24LongCat-Flash-ChatMeituan74
25LongCat-Flash-LiteMeituan73
26MAI-Code-1-FlashMicrosoft72
27Kimi K2 InstructMoonshot AI66
28Kimi K2-Instruct-0905Moonshot AI66
29Nemotron 3 Super (120B A12B)NVIDIA64
30o3OpenAI58
31Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team46
32Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team44
33Nemotron 3 Nano (30B A3B)NVIDIA42
34GPT-4oOpenAI24
35Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team13

35 of 35 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC