τ³-Banking Benchmark Leaderboard

AgenticLive

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

179

in our data

Data status

Live

Top score

33.4%

best on record

Top model

Kimi K3

Kimi

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A fintech customer-support benchmark from the τ-Knowledge framework that tests whether agents can navigate a large unstructured knowledge base and execute multi-step tool calls to resolve realistic banking workflows.

Leaderboard

Top 20 of 179 models we hold a score for.

1Kimi K3Kimi
33.4%2GPT-5.6 SolmaxOpenAI
33.0%3Claude Opus 5Adaptive Reasoning, High EffortAnthropic
32.8%4GPT-5.6 SolxhighOpenAI
32.6%5Grok 4.5highSpaceXAI
32.6%6GPT-5.6 TerramaxOpenAI
31.8%7Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
31.5%8GPT-5.5xhighOpenAI
31.3%9GPT-5.6 SolhighOpenAI
30.6%10Claude Sonnet 4.6Adaptive Reasoning, Max EffortAnthropic
30.5%11Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
30.3%12GPT-5.4xhighOpenAI
30.3%13GPT-5.5highOpenAI
29.5%14Claude Opus 4.7Adaptive Reasoning, Max EffortAnthropic
28.9%15Motif 3BetaMotif Technologies
28.9%16Claude Opus 5Adaptive Reasoning, Medium EffortAnthropic
28.7%17Claude Sonnet 5Adaptive Reasoning, Max EffortAnthropic
28.2%18JT-4.1 Flash 236B A21BChina Mobile
28.0%19Claude Opus 4.8Adaptive Reasoning, Max EffortAnthropic
27.6%20GPT-5.6 LunamaxOpenAI
27.2%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

A customer-support test set in banking: the agent has to find the right policy inside a large, unstructured knowledge base and then execute the multi-step tool calls that actually resolve the customer’s problem. Higher is better; the score is the share of cases carried through to a correct end state. Because it combines retrieval and action, a low score does not tell you which half broke — an agent can read the policy correctly and still fumble the transaction — so the ranking is more useful than the number. The customer and the knowledge base are both simulated, which makes this a measure of procedure-following rather than of how a model handles a genuinely distressed person.

© 2026 NYSGPT2525 LLC