Tau2 Airline

Evaluating Conversational Agents in a Dual-Control Environment

Models scored

23

evaluated

Modality

text

Category

communication

+2 more

Published

2025

arxiv.org

Citations

258

Semantic Scholar

Influential

40

citations

References

35

cited works

Venue

arXiv.org

published in

Abstract

Victor Barres, Honghua Dong, Soham Ray, Xujie Si, et al. (+1)

Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce $\tau^2$-bench, with four key contributions: 1) A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2) A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3) A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4) Fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, $\tau^2$-bench provides a controlled testbed for agents that must both reason effectively and guide user actions.

communicationreasoningtool calling

Search

#ModelLabScore
01LongCat-Flash-Thinking-2601Meituan77
02Nova 2 OmniAmazon69
03LongCat-Flash-ThinkingMeituan68
04GPT-5.1 ThinkingOpenAI67
05GPT-5.1OpenAI67
06GPT-5.1 InstantOpenAI67
07Nova 2 ProAmazon65
08Nova 2 LiteAmazon65
09o3OpenAI65
10Claude Haiku 4.5Anthropic64
11GPT-5OpenAI63
12Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team61
13Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team58
14LongCat-Flash-LiteMeituan58
15LongCat-Flash-ChatMeituan58
16Kimi K2 InstructMoonshot AI56
17Kimi K2-Instruct-0905Moonshot AI56
18Nemotron 3 Super (120B A12B)NVIDIA56
19Mercury 2Inception53
20Nemotron 3 Nano (30B A3B)NVIDIA48
21Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team46
22GPT-4oOpenAI46
23Qwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team44

23 of 23 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC