SciCode Benchmark Leaderboard

CodingLive

Category

Coding

Source

Artificial Analysis

evaluation of record

Models covered

545

in our data

Data status

Live

Top score

60.2%

best on record

Top model

Claude Fable 5

Anthropic

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A scientist-curated coding benchmark featuring 288 test set subproblems from 80 laboratory problems across 16 scientific disciplines.

Leaderboard

Top 20 of 545 models we hold a score for.

1Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
60.2%2Gemini 3.1 Pro PreviewGoogle
58.9%3Kimi K3Kimi
58.7%4Muse Spark 1.1xhighMeta
58.2%5GPT-5.6 SolhighOpenAI
56.9%6GPT-5.4xhighOpenAI
56.6%7GPT-5.6 SolmediumOpenAI
56.5%8GPT-5.5xhighOpenAI
56.1%9GPT-5.6 SolmaxOpenAI
56.1%10Gemini 3 Pro PreviewhighGoogle
56.1%11GPT-5.6 SolxhighOpenAI
56.0%12GPT-5.5highOpenAI
55.9%13Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
55.7%14GPT-5.6 SollowOpenAI
55.4%15Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
55.0%16GPT-5.2 CodexxhighOpenAI
54.6%17Claude Opus 4.7Adaptive Reasoning, Max EffortAnthropic
54.5%18Claude Opus 5Adaptive Reasoning, High EffortAnthropic
54.3%19Grok 4.5highSpaceXAI
54.1%20GPT-5.6 TerramaxOpenAI
53.9%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Real scientific programming: 80 research problems drawn from 16 fields, broken into 288 subproblems, each written by a working scientist and marked by running the code. Higher is better, and scores sit far below general coding benchmarks because the problems require domain knowledge as well as programming ability. Subproblems build on one another, so a model can solve most of the pieces and still fail the parent problem — subproblem and problem-level scores tell noticeably different stories. Passing the tests means the code ran correctly, not that a scientist would accept the approach it took.

© 2026 NYSGPT2525 LLC