τ²-Bench Telecom Benchmark Leaderboard

AgenticLive

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

440

in our data

Data status

Live

Top score

99.1%

best on record

Top model

GLM-5.2

Z AI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A dual-control conversational AI benchmark simulating technical support scenarios where both agent and user must coordinate actions to resolve telecom service issues.

Leaderboard

Top 20 of 440 models we hold a score for.

1GLM-5.2maxZ AI
99.1%2JT-35B-FlashChina Mobile
99.1%3GLM-4.7-FlashReasoningZ AI
98.8%4Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
98.5%5GLM 5V TurboReasoningZ AI
98.5%6GLM-5-TurboZ AI
98.5%7Step 3.7 FlashStepFun
98.5%8GLM-5ReasoningZ AI
98.2%9GLM-5.1ReasoningZ AI
97.7%10Grok 4.3highSpaceXAI
97.7%11Qwen3.6 PlusAlibaba
97.7%12GLM-5Non-reasoningZ AI
97.4%13GLM-5.1Non-reasoningZ AI
97.1%14Grok 4.20 0309ReasoningSpaceXAI
96.5%15DeepSeek V4 ProReasoning, Max EffortDeepSeek
96.2%16GLM-4.7ReasoningZ AI
95.9%17Kimi K2.5ReasoningKimi
95.9%18Kimi K2.6Kimi
95.9%19Qwen3.6 Max PreviewAlibaba
95.9%20DeepSeek V4 FlashReasoning, High EffortDeepSeek
95.6%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

A telecom support scenario in which the model and the simulated customer each control part of the system, so neither can finish alone — the agent has to get the user to do something, then act on whatever happened. Higher is better; the score is the share of scenarios brought to a correct end state. The dual-control design is what makes it hard: models that are strong at single-turn tool use often fail here because they issue instructions the user cannot actually carry out. The user is played by another language model, so some part of every score reflects the simulator’s behavior rather than the agent under test.

© 2026 NYSGPT2525 LLC