Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
440
in our data
Data status
Live
Top score
99.1%
best on record
Top model
GLM-5.2
Z AI
Updated
2026-07-30
last ingest
A dual-control conversational AI benchmark simulating technical support scenarios where both agent and user must coordinate actions to resolve telecom service issues.
Leaderboard
Top 20 of 440 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
A telecom support scenario in which the model and the simulated customer each control part of the system, so neither can finish alone — the agent has to get the user to do something, then act on whatever happened. Higher is better; the score is the share of scenarios brought to a correct end state. The dual-control design is what makes it hard: models that are strong at single-turn tool use often fail here because they issue instructions the user cannot actually carry out. The user is played by another language model, so some part of every score reflects the simulator’s behavior rather than the agent under test.