Category
Composite
Source
Artificial Analysis
evaluation of record
Models covered
573
in our data
Data status
Live
Top score
60.7
best on record
Top model
Claude Opus 5
Anthropic
Updated
2026-07-30
last ingest
A composite benchmark aggregating nine challenging evaluations to provide a holistic measure of AI capabilities across mathematics, science, coding, and reasoning.
Leaderboard
Top 20 of 573 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
A single 0–100 score that averages a model’s results across nine harder evaluations, so one number stands in for reasoning, science, coding, and math at once. Higher is better: the frontier currently sits in the low 60s, and most models in production land between 20 and 50. Because it is an average, a 10-point gap usually means a model is broadly stronger rather than better at any one thing — and a model can lift its index by fixing the two or three components it is worst at. The composite is only as current as its component set: when a saturated evaluation is swapped out, scores move across the whole board, so compare indexes within a version rather than across time.