TAU-bench Retail

A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Models scored

25

evaluated

Modality

text

Category

communication

+2 more

Published

2024

arxiv.org

Citations

636

Semantic Scholar

Influential

93

citations

References

31

cited works

Venue

arXiv.org

published in

Abstract

Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan

Existing benchmarks do not test language agents on their interaction with human users or ability to follow domain-specific rules, both of which are vital for deploying them in real world applications. We propose $\tau$-bench, a benchmark emulating dynamic conversations between a user (simulated by language models) and a language agent provided with domain-specific API tools and policy guidelines. We employ an efficient and faithful evaluation process that compares the database state at the end of a conversation with the annotated goal state. We also propose a new metric (pass^k) to evaluate the reliability of agent behavior over multiple trials. Our experiments show that even state-of-the-art function calling agents (like gpt-4o) succeed on<50% of the tasks, and are quite inconsistent (pass^8<25% in retail). Our findings point to the need for methods that can improve the ability of agents to act consistently and follow rules reliably.

communicationreasoningtool calling

Search

#ModelLabScore
01Claude Sonnet 4.5Anthropic86
02Claude Opus 4.1Anthropic82
03Claude Opus 4Anthropic81
04Claude 3.7 SonnetAnthropic81
05Claude Sonnet 4Anthropic81
06GLM-4.5Zhipu AI80
07GLM-4.5-AirZhipu AI78
08Qwen3-Coder 480B A35B InstructAlibaba Cloud / Qwen Team78
09o4-miniOpenAI72
10o1OpenAI71
11Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team70
12Claude 3.5 SonnetAnthropic69
13GPT-4.5OpenAI68
14GPT-4.1OpenAI68
15Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team68
16GPT OSS 120BOpenAI68
17MiniMax M1 40KMiniMax68
18MiniMax M1 80KMiniMax64
19Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team61
20GPT-4oOpenAI60
21o3-miniOpenAI58
22GPT-4.1 miniOpenAI56
23GPT OSS 20BOpenAI55
24Claude 3.5 HaikuAnthropic51
25GPT-4.1 nanoOpenAI23

25 of 25 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC