LiveCodeBench Benchmark Leaderboard

CodingLive

Category

Coding

Source

Artificial Analysis

evaluation of record

Models covered

343

in our data

Data status

Live

Top score

91.7%

best on record

Top model

Gemini 3 Pro Preview

Google

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A contamination-free coding benchmark that continuously harvests fresh competitive programming problems from LeetCode, AtCoder, and CodeForces, evaluating code generation, self-repair, and execution.

Leaderboard

Top 20 of 343 models we hold a score for.

1Gemini 3 Pro PreviewhighGoogle
91.7%2Gemini 3 Flash PreviewReasoningGoogle
90.8%3DeepSeek V3.2 SpecialeDeepSeek
89.6%4GLM-4.7ReasoningZ AI
89.4%5GPT-5.2mediumOpenAI
89.4%6GPT-5.2xhighOpenAI
88.9%7gpt-oss-120bhighOpenAI
87.8%8Claude Opus 4.5ReasoningAnthropic
87.1%9GPT-5.1highOpenAI
86.8%10MiMo-V2-FlashReasoningXiaomi
86.8%11DeepSeek V3.2ReasoningDeepSeek
86.2%12o4-minihighOpenAI
85.9%13Gemini 3 Pro PreviewlowGoogle
85.7%14Kimi K2 ThinkingKimi
85.3%15GPT-5.1 CodexhighOpenAI
84.9%16GPT-5highOpenAI
84.6%17GPT-5 CodexhighOpenAI
84.0%18GPT-5 minihighOpenAI
83.8%19GPT-5.1 Codex minihighOpenAI
83.6%20Grok 4 FastReasoningSpaceXAI
83.2%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Competitive programming problems pulled continuously from LeetCode, AtCoder, and Codeforces, marked by running the code, and filtered by date so a model can be tested only on problems published after its training cut-off. Higher is better, and the date filtering is the entire design — it is the main defense against a model having simply memorized the answers. Because the problem pool keeps growing, a score only means something with its date window attached, and two published LiveCodeBench figures are frequently not the same test. It measures algorithmic puzzle-solving, which correlates with but is not the same thing as writing software someone else has to maintain.

© 2026 NYSGPT2525 LLC