Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
—
none ingested
Data status
Demo Purposes Only
invented numbers
Top score
—
not measured
Top model
—
Updated
—
no ingest
Artificial Analysis' implementation of IBM's ITBench benchmark, testing AI agents on Kubernetes incident root-cause analysis from offline incident snapshots. The agent inspects alerts, events, traces, and topology and identifies the contributing-factor entities (deployments, pods, namespaces, network policies, etc.) responsible for the failure.
Leaderboard
The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.
Demo purposes only · invented numbers · not a measurement
Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Gives an agent a frozen snapshot of a Kubernetes incident — alerts, events, traces, topology — and asks the first question any on-call engineer asks: which components actually caused this? Scoring compares the entities the agent named against the known contributing factors, so higher is better, and naming too many hurts as much as missing one. Because the snapshots are offline, the agent cannot run a command and watch what changes, which removes the most useful debugging tool a real engineer has and makes this harder than the live job in one specific way. It is an independent implementation of IBM’s benchmark, so compare within this leaderboard rather than against IBM’s published figures.