TAU-bench Airline

A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Models scored

23

evaluated

Modality

text

Category

communication

+2 more

Published

2024

arxiv.org

Citations

636

Semantic Scholar

Influential

93

citations

References

31

cited works

Venue

arXiv.org

published in

Abstract

Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan

Existing benchmarks do not test language agents on their interaction with human users or ability to follow domain-specific rules, both of which are vital for deploying them in real world applications. We propose $\tau$-bench, a benchmark emulating dynamic conversations between a user (simulated by language models) and a language agent provided with domain-specific API tools and policy guidelines. We employ an efficient and faithful evaluation process that compares the database state at the end of a conversation with the annotated goal state. We also propose a new metric (pass^k) to evaluate the reliability of agent behavior over multiple trials. Our experiments show that even state-of-the-art function calling agents (like gpt-4o) succeed on<50% of the tasks, and are quite inconsistent (pass^8<25% in retail). Our findings point to the need for methods that can improve the ability of agents to act consistently and follow rules reliably.

communicationreasoningtool calling

Search

#ModelLabScore
01Claude Sonnet 4.5Anthropic70
02MiniMax M1 80KMiniMax62
03GLM-4.5-AirZhipu AI61
04GLM-4.5Zhipu AI60
05MiniMax M1 40KMiniMax60
06Qwen3-Coder 480B A35B InstructAlibaba Cloud / Qwen Team60
07Claude Sonnet 4Anthropic60
08Claude Opus 4Anthropic60
09Claude 3.7 SonnetAnthropic58
10Claude Opus 4.1Anthropic56
11GPT-4.5OpenAI50
12o1OpenAI50
13GPT-4.1OpenAI49
14o4-miniOpenAI49
15Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team49
16Claude 3.5 SonnetAnthropic46
17Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team46
18Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team44
19GPT-4oOpenAI43
20GPT-4.1 miniOpenAI36
21o3-miniOpenAI32
22Claude 3.5 HaikuAnthropic23
23GPT-4.1 nanoOpenAI14

23 of 23 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC