Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
—
none ingested
Data status
Demo Purposes Only
invented numbers
Top score
—
not measured
Top model
—
Updated
—
no ingest
Artificial Analysis' independent implementation of ServiceNow's EnterpriseOps-Gym, an agentic benchmark testing whether LLM agents can complete stateful, multi-step enterprise workflows across eight business domains via live tool use, graded on the final state of the underlying databases.
Leaderboard
The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.
Demo purposes only · invented numbers · not a measurement
Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Tests whether an agent can complete multi-step enterprise workflows across eight business domains by actually calling tools, then grades the state the databases were left in rather than what the agent said it did. Higher is better, and grading on final state is unusually strict: an agent that narrates a correct plan but writes the wrong record scores nothing for that task. That makes it one of the more honest agentic measurements available, and also one of the lowest-scoring. It is an independent implementation of ServiceNow’s benchmark, so compare within this leaderboard rather than against ServiceNow’s own published figures.