Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
179
in our data
Data status
Live
Top score
33.4%
best on record
Top model
Kimi K3
Kimi
Updated
2026-07-30
last ingest
A fintech customer-support benchmark from the τ-Knowledge framework that tests whether agents can navigate a large unstructured knowledge base and execute multi-step tool calls to resolve realistic banking workflows.
Leaderboard
Top 20 of 179 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
A customer-support test set in banking: the agent has to find the right policy inside a large, unstructured knowledge base and then execute the multi-step tool calls that actually resolve the customer’s problem. Higher is better; the score is the share of cases carried through to a correct end state. Because it combines retrieval and action, a low score does not tell you which half broke — an agent can read the policy correctly and still fumble the transaction — so the ranking is more useful than the number. The customer and the knowledge base are both simulated, which makes this a measure of procedure-following rather than of how a model handles a genuinely distressed person.