ITBench-AA Benchmark Leaderboard

AgenticDemo Purposes Only

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

none ingested

Data status

Demo Purposes Only

invented numbers

Top score

not measured

Top model

Updated

no ingest

Full results

artificialanalysis.ai

on the source

Artificial Analysis' implementation of IBM's ITBench benchmark, testing AI agents on Kubernetes incident root-cause analysis from offline incident snapshots. The agent inspects alerts, events, traces, and topology and identifies the contributing-factor entities (deployments, pods, namespaces, network policies, etc.) responsible for the failure.

Leaderboard

The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.

Demo purposes only · invented numbers · not a measurement

1Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
44.0%2GPT-5.6 SolmaxOpenAI
40.7%3Kimi K3Kimi
36.7%4Grok 4.5highSpaceXAI
33.2%5GLM-5.2maxZ AI
30.7%6Muse Spark 1.1xhighMeta
28.3%7Gemini 3.5 FlashhighGoogle
24.7%8Qwen3.7 MaxAlibaba
22.7%9MiniMax-M3MiniMax
20.2%10DeepSeek V4 ProReasoning, Max EffortDeepSeek
18.1%11Motif 3BetaMotif Technologies
16.4%12MiMo-V2.5-ProXiaomi
15.2%13Hy3Tencent
13.4%14Nex-N2-ProNex AGI
12.4%15InklingxhighThinking Machines
10.8%16Agnes 2.5 Pro AlphaSapiens AI
9.6%17JT-4.1 Flash 236B A21BChina Mobile
8.4%18Nemotron 3 Ultra 550B A55BReasoningNVIDIA
7.7%19Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
6.9%20Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
6.3%

Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Gives an agent a frozen snapshot of a Kubernetes incident — alerts, events, traces, topology — and asks the first question any on-call engineer asks: which components actually caused this? Scoring compares the entities the agent named against the known contributing factors, so higher is better, and naming too many hurts as much as missing one. Because the snapshots are offline, the agent cannot run a command and watch what changes, which removes the most useful debugging tool a real engineer has and makes this harder than the live job in one specific way. It is an independent implementation of IBM’s benchmark, so compare within this leaderboard rather than against IBM’s published figures.

© 2026 NYSGPT2525 LLC