EnterpriseOps-Gym-AA Benchmark Leaderboard

AgenticDemo Purposes Only

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

none ingested

Data status

Demo Purposes Only

invented numbers

Top score

not measured

Top model

Updated

no ingest

Full results

artificialanalysis.ai

on the source

Artificial Analysis' independent implementation of ServiceNow's EnterpriseOps-Gym, an agentic benchmark testing whether LLM agents can complete stateful, multi-step enterprise workflows across eight business domains via live tool use, graded on the final state of the underlying databases.

Leaderboard

The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.

Demo purposes only · invented numbers · not a measurement

1Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
34.0%2GPT-5.6 SolmaxOpenAI
31.1%3Kimi K3Kimi
27.9%4Grok 4.5highSpaceXAI
25.3%5GLM-5.2maxZ AI
23.1%6Muse Spark 1.1xhighMeta
20.3%7Gemini 3.5 FlashhighGoogle
18.7%8Qwen3.7 MaxAlibaba
17.3%9MiniMax-M3MiniMax
15.1%10DeepSeek V4 ProReasoning, Max EffortDeepSeek
13.9%11Motif 3BetaMotif Technologies
12.2%12MiMo-V2.5-ProXiaomi
10.7%13Hy3Tencent
9.5%14Nex-N2-ProNex AGI
8.7%15InklingxhighThinking Machines
7.6%16Agnes 2.5 Pro AlphaSapiens AI
6.8%17JT-4.1 Flash 236B A21BChina Mobile
6.0%18Nemotron 3 Ultra 550B A55BReasoningNVIDIA
5.4%19Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
5.0%20Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
4.4%

Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Tests whether an agent can complete multi-step enterprise workflows across eight business domains by actually calling tools, then grades the state the databases were left in rather than what the agent said it did. Higher is better, and grading on final state is unusually strict: an agent that narrates a correct plan but writes the wrong record scores nothing for that task. That makes it one of the more honest agentic measurements available, and also one of the lowest-scoring. It is an independent implementation of ServiceNow’s benchmark, so compare within this leaderboard rather than against ServiceNow’s own published figures.

© 2026 NYSGPT2525 LLC