Category
Coding
Source
Artificial Analysis
evaluation of record
Models covered
545
in our data
Data status
Live
Top score
60.2%
best on record
Top model
Claude Fable 5
Anthropic
Updated
2026-07-30
last ingest
A scientist-curated coding benchmark featuring 288 test set subproblems from 80 laboratory problems across 16 scientific disciplines.
Leaderboard
Top 20 of 545 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Real scientific programming: 80 research problems drawn from 16 fields, broken into 288 subproblems, each written by a working scientist and marked by running the code. Higher is better, and scores sit far below general coding benchmarks because the problems require domain knowledge as well as programming ability. Subproblems build on one another, so a model can solve most of the pieces and still fail the parent problem — subproblem and problem-level scores tell noticeably different stories. Passing the tests means the code ran correctly, not that a scientist would accept the approach it took.