LiveCodeBench

Holistic and Contamination Free Evaluation of Large Language Models for Code

Models scored

73

evaluated

Modality

text

Category

code

+2 more

Published

2024

arxiv.org

Citations

1,637

Semantic Scholar

Influential

263

citations

References

90

cited works

Venue

International Conference on Learning Representations

published in

Abstract

Naman Jain, King Han, Alex Gu, Wen-Ding Li, et al. (+6)

Large Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from both academia and industry. However, as new and improved LLMs are developed, existing evaluation benchmarks (e.g., HumanEval, MBPP) are no longer sufficient for assessing their capabilities. In this work, we propose LiveCodeBench, a comprehensive and contamination-free evaluation of LLMs for code, which continuously collects new problems over time from contests across three competition platforms, namely LeetCode, AtCoder, and CodeForces. Notably, our benchmark also focuses on a broader range of code related capabilities, such as self-repair, code execution, and test output prediction, beyond just code generation. Currently, LiveCodeBench hosts four hundred high-quality coding problems that were published between May 2023 and May 2024. We have evaluated 18 base LLMs and 34 instruction-tuned LLMs on LiveCodeBench. We present empirical findings on contamination, holistic performance comparisons, potential overfitting in existing benchmarks as well as individual model comparisons. We will release all prompts and model completions for further community analysis, along with a general toolkit for adding new scenarios and model

codegeneralreasoning

Search

#ModelLabScore
01DeepSeek-V4-Pro-MaxDeepSeek94
02DeepSeek-V4-Flash-MaxDeepSeek92
03DeepSeek-V3.2DeepSeek83
04DeepSeek-V3.2 (Thinking)DeepSeek83
05MiniMax M2MiniMax83
06LongCat-Flash-Thinking-2601Meituan83
07Nemotron 3 Super (120B A12B)NVIDIA81
08Grok-3 MinixAI80
09Grok 4 FastxAI80
10Grok-4 HeavyxAI79
11Grok-3xAI79
12LongCat-Flash-ThinkingMeituan79
13Grok-4xAI79
14MiniMax M2.1MiniMax78
15Nova 2 ProAmazon75
16DeepSeek-V3.2-ExpDeepSeek74
17DeepSeek-R1-0528DeepSeek73
18GLM-4.5Zhipu AI73
19Nemotron Nano 9B v2NVIDIA71
20Nova 2 LiteAmazon71
21GLM-4.5-AirZhipu AI71
22Qwen3 235B A22BAlibaba Cloud / Qwen Team71
23Gemini 2.5 Pro Preview 06-05Google69
24Mercury 2Inception67
25Llama 3.1 Nemotron Ultra 253B v1NVIDIA66
26Qwen3 32BAlibaba Cloud / Qwen Team66
27MiniMax M1 80KMiniMax65
28Ministral 3 (14B Reasoning 2512)Mistral AI65
29Mistral Small 4Mistral AI64
30QwQ-32BAlibaba Cloud / Qwen Team63
31Qwen3 30B A3BAlibaba Cloud / Qwen Team63
32MiniMax M1 40KMiniMax62
33Ministral 3 (8B Reasoning 2512)Mistral AI62
34DeepSeek R1 Distill Llama 70BDeepSeek57
35DeepSeek R1 Distill Qwen 32BDeepSeek57
36DeepSeek-V3.1DeepSeek56
37Qwen2.5 72B InstructAlibaba Cloud / Qwen Team56
38Min istral 3 (3B Reasoning 2512)Mistral AI55
39Phi 4 ReasoningMicrosoft54
40Kimi K2-Instruct-0905Moonshot AI54
41DeepSeek R1 Distill Qwen 14BDeepSeek53
42Phi 4 Reasoning PlusMicrosoft53
43Magistral Small 2506Mistral AI51
44Magistral MediumMistral AI50
45QwQ-32B-PreviewAlibaba Cloud / Qwen Team50
46DeepSeek R1 ZeroDeepSeek50
47DeepSeek-V3 0324DeepSeek49
48LongCat-Flash-ChatMeituan48
49Llama 4 MaverickMeta43
50DeepSeek R1 Distill Llama 8BDeepSeek40
51DeepSeek R1 Distill Qwen 7BDeepSeek38
52DeepSeek-V3DeepSeek38
53Gemini 2.0 FlashGoogle35
54Mistral Large 3 (675B Base)Mistral AI34
55Mistral Large 3 (675B Instruct 2512 NVFP4)Mistral AI34
56Mistral Large 3 (675B Instruct 2512)Mistral AI34
57Mistral Large 3 (675B Instruct 2512 Eagle)Mistral AI34
58Gemini 2.5 Flash-LiteGoogle34
59Llama 4 ScoutMeta33
60Qwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team31

60 of 73 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC