Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
7
in our data
Data status
Live
Top score
37.6%
best on record
Top model
Kimi K3
Moonshot AI
Updated
2026-07-30
last ingest
Artificial Analysis' implementation of the APEX-Agents benchmark, testing AI agents on long-horizon, cross-application tasks in professional-services environments with realistic application tooling.
Leaderboard
Top 7 of 7 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Measures whether an agent can carry a professional-services task across several applications without losing the thread — work that starts in email, moves through a document, and ends in a system of record. Higher is better; scores are the share of task objectives completed, and current numbers are low, with leaders finishing roughly a third. Long, cross-application work compounds errors, so one wrong step early can zero out an otherwise capable run, which makes these scores noisier than knowledge tests. This is an independent implementation, so the numbers are comparable within this leaderboard but not against figures published by the benchmark’s original authors.