LiveBench

A Challenging, Contamination-Limited LLM Benchmark

Models scored

38

evaluated

Modality

text

Category

general

+2 more

Published

2024

arxiv.org

Citations

153

Semantic Scholar

Influential

11

citations

References

69

cited works

Venue

International Conference on Learning Representations

published in

Abstract

Colin White, Samuel Dooley, Manley Roberts, Arka Pal, et al. (+14)

Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models.

generalmathreasoning

Search

#ModelLabScore
01o3-miniOpenAI85
02GPT-5.5OpenAI81
03GPT-5.4OpenAI80
04Gemini 3.1 ProGoogle80
05Claude Fable 5Anthropic78
06Claude Opus 4.8Anthropic77
07Qwen3 235B A22BAlibaba Cloud / Qwen Team77
08Claude Opus 4.7Anthropic77
09Kimi K2 InstructMoonshot AI76
10Kimi K2-Instruct-0905Moonshot AI76
11Claude Opus 4.6Anthropic76
12Claude Opus 4.5Anthropic76
13Claude Sonnet 4.6Anthropic75
14Gemini 3.5 FlashGoogle75
15Qwen3 32BAlibaba Cloud / Qwen Team75
16GPT-5.2OpenAI75
17Qwen3 30B A3BAlibaba Cloud / Qwen Team74
18GPT-5.2 CodexOpenAI74
19Qwen3.7 MaxAlibaba Cloud / Qwen Team74
20DeepSeek-V4-Pro-MaxDeepSeek74
21Gemini 3 ProGoogle73
22QwQ-32BAlibaba Cloud / Qwen Team73
23GPT-5.3 CodexOpenAI73
24Gemini 3 FlashGoogle72
25Kimi K2.6Moonshot AI72
26GPT-5.1 HighOpenAI72
27Kimi K2.7 CodeMoonshot AI72
28Qwen3.6 PlusAlibaba Cloud / Qwen Team71
29GLM-5.1Zhipu AI70
30GPT-5.4 nanoOpenAI70
31MiniMax M3MiniMax70
32Kimi K2.5Moonshot AI69
33o1OpenAI67
34o1-previewOpenAI52
35Qwen2.5 72B InstructAlibaba Cloud / Qwen Team52
36Phi 4Microsoft48
37Qwen2.5 7B InstructAlibaba Cloud / Qwen Team36
38Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team30

38 of 38 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC