LiveCodeBench v6
Holistic and Contamination Free Evaluation of Large Language Models for Code
Models scored
53
evaluated
Modality
text
Category
general
+1 more
Published
2024
arxiv.org
Citations
1,637
Semantic Scholar
Influential
263
citations
References
90
cited works
Venue
International Conference on Learning Representations
published in
Abstract
Naman Jain, King Han, Alex Gu, Wen-Ding Li, et al. (+6)
Large Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from both academia and industry. However, as new and improved LLMs are developed, existing evaluation benchmarks (e.g., HumanEval, MBPP) are no longer sufficient for assessing their capabilities. In this work, we propose LiveCodeBench, a comprehensive and contamination-free evaluation of LLMs for code, which continuously collects new problems over time from contests across three competition platforms, namely LeetCode, AtCoder, and CodeForces. Notably, our benchmark also focuses on a broader range of code related capabilities, such as self-repair, code execution, and test output prediction, beyond just code generation. Currently, LiveCodeBench hosts four hundred high-quality coding problems that were published between May 2023 and May 2024. We have evaluated 18 base LLMs and 34 instruction-tuned LLMs on LiveCodeBench. We present empirical findings on contamination, holistic performance comparisons, potential overfitting in existing benchmarks as well as individual model comparisons. We will release all prompts and model completions for further community analysis, along with a general toolkit for adding new scenarios and model
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Qwen3.7 Max | Alibaba Cloud / Qwen Team | 92 |
| 02 | Kimi K2.6 | Moonshot AI | 90 |
| 03 | Qwen3.7-Plus | Alibaba Cloud / Qwen Team | 90 |
| 04 | Nemotron 3 Ultra (550B A55B) | NVIDIA | 89 |
| 05 | Seed 2.0 Pro | ByteDance | 88 |
| 06 | MAI-Thinking-1 | Microsoft | 88 |
| 07 | Qwen3.6 Plus | Alibaba Cloud / Qwen Team | 87 |
| 08 | Step-3.5-Flash | StepFun | 86 |
| 09 | Kimi K2.5 | Moonshot AI | 85 |
| 10 | GLM-4.7 | Zhipu AI | 85 |
| 11 | Qwen3.6-27B | Alibaba Cloud / Qwen Team | 84 |
| 12 | Qwen3.5-397B-A17B | Alibaba Cloud / Qwen Team | 84 |
| 13 | Kimi K2-Thinking-0905 | Moonshot AI | 83 |
| 14 | GLM-4.6 | Zhipu AI | 83 |
| 15 | GPT OSS 120B High | OpenAI | 82 |
| 16 | Seed 2.0 Lite | ByteDance | 82 |
| 17 | K-EXAONE-236B-A23B | LG AI Research | 81 |
| 18 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 81 |
| 19 | MiMo-V2-Flash | Xiaomi | 81 |
| 20 | Qwen3.6-35B-A3B | Alibaba Cloud / Qwen Team | 80 |
| 21 | Gemma 4 31B | 80 | |
| 22 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 79 |
| 23 | Gemma 4 26B-A4B | 77 | |
| 24 | Qwen3.5-35B-A3B | Alibaba Cloud / Qwen Team | 75 |
| 25 | Qwen3-235B-A22B-Thinking-2507 | Alibaba Cloud / Qwen Team | 74 |
| 26 | Gemma 4 12B | 72 | |
| 27 | Sarvam-105B | Sarvam AI | 72 |
| 28 | North Mini Code 1.0 | Cohere | 70 |
| 29 | Qwen3 VL 235B A22B Thinking | Alibaba Cloud / Qwen Team | 70 |
| 30 | Sarvam-30B | Sarvam AI | 70 |
| 31 | DiffusionGemma 26B-A4B | 69 | |
| 32 | Qwen3 Max | Alibaba Cloud / Qwen Team | 69 |
| 33 | Qwen3-Next-80B-A3B-Thinking | Alibaba Cloud / Qwen Team | 69 |
| 34 | Nemotron 3 Nano (30B A3B) | NVIDIA | 68 |
| 35 | Qwen3.5-9B | Alibaba Cloud / Qwen Team | 66 |
| 36 | Qwen3 VL 32B Thinking | Alibaba Cloud / Qwen Team | 66 |
| 37 | Qwen3 VL 30B A3B Thinking | Alibaba Cloud / Qwen Team | 64 |
| 38 | Qwen3 VL 8B Thinking | Alibaba Cloud / Qwen Team | 59 |
| 39 | Qwen3-Next-80B-A3B-Instruct | Alibaba Cloud / Qwen Team | 57 |
| 40 | Qwen3.5-4B | Alibaba Cloud / Qwen Team | 56 |
| 41 | Qwen3 VL 235B A22B Instruct | Alibaba Cloud / Qwen Team | 54 |
| 42 | Kimi K2 Instruct | Moonshot AI | 54 |
| 43 | MiniCPM-SALA | OpenBMB | 52 |
| 44 | Gemma 4 E4B | 52 | |
| 45 | Qwen3-235B-A22B-Instruct-2507 | Alibaba Cloud / Qwen Team | 52 |
| 46 | Qwen3 VL 4B Thinking | Alibaba Cloud / Qwen Team | 51 |
| 47 | Gemma 4 E2B | 44 | |
| 48 | Qwen3 VL 32B Instruct | Alibaba Cloud / Qwen Team | 44 |
| 49 | Qwen3 VL 30B A3B Instruct | Alibaba Cloud / Qwen Team | 43 |
| 50 | MiMo-V2.5-Pro | Xiaomi | 40 |
| 51 | Qwen3 VL 8B Instruct | Alibaba Cloud / Qwen Team | 39 |
| 52 | Qwen3 VL 4B Instruct | Alibaba Cloud / Qwen Team | 38 |
| 53 | Kimi K2 Base | Moonshot AI | 26 |
53 of 53 models · score normalized 0–100 where available