Category
Coding
Source
Artificial Analysis
evaluation of record
Models covered
343
in our data
Data status
Live
Top score
91.7%
best on record
Top model
Gemini 3 Pro Preview
Updated
2026-07-30
last ingest
A contamination-free coding benchmark that continuously harvests fresh competitive programming problems from LeetCode, AtCoder, and CodeForces, evaluating code generation, self-repair, and execution.
Leaderboard
Top 20 of 343 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Competitive programming problems pulled continuously from LeetCode, AtCoder, and Codeforces, marked by running the code, and filtered by date so a model can be tested only on problems published after its training cut-off. Higher is better, and the date filtering is the entire design — it is the main defense against a model having simply memorized the answers. Because the problem pool keeps growing, a score only means something with its date window attached, and two published LiveCodeBench figures are frequently not the same test. It measures algorithmic puzzle-solving, which correlates with but is not the same thing as writing software someone else has to maintain.