APEX-Agents-AA Benchmark Leaderboard

AgenticLive

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

7

in our data

Data status

Live

Top score

37.6%

best on record

Top model

Kimi K3

Moonshot AI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

Artificial Analysis' implementation of the APEX-Agents benchmark, testing AI agents on long-horizon, cross-application tasks in professional-services environments with realistic application tooling.

Leaderboard

Top 7 of 7 models we hold a score for.

1Kimi K3Moonshot AI
37.6%2Seed 2.1 ProByteDance
33.8%3Gemini 3.1 ProGoogle
33.5%4Seed 2.1 TurboByteDance
29.2%5Kimi K2.6Moonshot AI
27.9%6MiniMax M3MiniMax
27.7%7Hy3Tencent
25.6%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Measures whether an agent can carry a professional-services task across several applications without losing the thread — work that starts in email, moves through a document, and ends in a system of record. Higher is better; scores are the share of task objectives completed, and current numbers are low, with leaders finishing roughly a third. Long, cross-application work compounds errors, so one wrong step early can zero out an otherwise capable run, which makes these scores noisier than knowledge tests. This is an independent implementation, so the numbers are comparable within this leaderboard but not against figures published by the benchmark’s original authors.

© 2026 NYSGPT2525 LLC