LiveCodeBench v6

Holistic and Contamination Free Evaluation of Large Language Models for Code

Models scored

53

evaluated

Modality

text

Category

general

+1 more

Published

2024

arxiv.org

Citations

1,637

Semantic Scholar

Influential

263

citations

References

90

cited works

Venue

International Conference on Learning Representations

published in

Abstract

Naman Jain, King Han, Alex Gu, Wen-Ding Li, et al. (+6)

Large Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from both academia and industry. However, as new and improved LLMs are developed, existing evaluation benchmarks (e.g., HumanEval, MBPP) are no longer sufficient for assessing their capabilities. In this work, we propose LiveCodeBench, a comprehensive and contamination-free evaluation of LLMs for code, which continuously collects new problems over time from contests across three competition platforms, namely LeetCode, AtCoder, and CodeForces. Notably, our benchmark also focuses on a broader range of code related capabilities, such as self-repair, code execution, and test output prediction, beyond just code generation. Currently, LiveCodeBench hosts four hundred high-quality coding problems that were published between May 2023 and May 2024. We have evaluated 18 base LLMs and 34 instruction-tuned LLMs on LiveCodeBench. We present empirical findings on contamination, holistic performance comparisons, potential overfitting in existing benchmarks as well as individual model comparisons. We will release all prompts and model completions for further community analysis, along with a general toolkit for adding new scenarios and model

generalreasoning

Search

#ModelLabScore
01Qwen3.7 MaxAlibaba Cloud / Qwen Team92
02Kimi K2.6Moonshot AI90
03Qwen3.7-PlusAlibaba Cloud / Qwen Team90
04Nemotron 3 Ultra (550B A55B)NVIDIA89
05Seed 2.0 ProByteDance88
06MAI-Thinking-1Microsoft88
07Qwen3.6 PlusAlibaba Cloud / Qwen Team87
08Step-3.5-FlashStepFun86
09Kimi K2.5Moonshot AI85
10GLM-4.7Zhipu AI85
11Qwen3.6-27BAlibaba Cloud / Qwen Team84
12Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team84
13Kimi K2-Thinking-0905Moonshot AI83
14GLM-4.6Zhipu AI83
15GPT OSS 120B HighOpenAI82
16Seed 2.0 LiteByteDance82
17K-EXAONE-236B-A23BLG AI Research81
18Qwen3.5-27BAlibaba Cloud / Qwen Team81
19MiMo-V2-FlashXiaomi81
20Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team80
21Gemma 4 31BGoogle80
22Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team79
23Gemma 4 26B-A4BGoogle77
24Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team75
25Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team74
26Gemma 4 12BGoogle72
27Sarvam-105BSarvam AI72
28North Mini Code 1.0Cohere70
29Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team70
30Sarvam-30BSarvam AI70
31DiffusionGemma 26B-A4BGoogle69
32Qwen3 MaxAlibaba Cloud / Qwen Team69
33Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team69
34Nemotron 3 Nano (30B A3B)NVIDIA68
35Qwen3.5-9BAlibaba Cloud / Qwen Team66
36Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team66
37Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team64
38Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team59
39Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team57
40Qwen3.5-4BAlibaba Cloud / Qwen Team56
41Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team54
42Kimi K2 InstructMoonshot AI54
43MiniCPM-SALAOpenBMB52
44Gemma 4 E4BGoogle52
45Qwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team52
46Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team51
47Gemma 4 E2BGoogle44
48Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team44
49Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team43
50MiMo-V2.5-ProXiaomi40
51Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team39
52Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team38
53Kimi K2 BaseMoonshot AI26

53 of 53 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC