HumanEval

Evaluating Large Language Models Trained on Code

Models scored

66

evaluated

Modality

text

Category

code

+1 more

Published

2021

arxiv.org

Citations

9,942

Semantic Scholar

Influential

1,534

citations

References

127

cited works

Venue

arXiv.org

published in

Abstract

Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, et al. (+49)

We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities. A distinct production version of Codex powers GitHub Copilot. On HumanEval, a new evaluation set we release to measure functional correctness for synthesizing programs from docstrings, our model solves 28.8% of the problems, while GPT-3 solves 0% and GPT-J solves 11.4%. Furthermore, we find that repeated sampling from the model is a surprisingly effective strategy for producing working solutions to difficult prompts. Using this method, we solve 70.2% of our problems with 100 samples per problem. Careful investigation of our model reveals its limitations, including difficulty with docstrings describing long chains of operations and with binding operations to variables. Finally, we discuss the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics.

Search

#ModelLabScore
01MiniCPM-SALAOpenBMB95
02Kimi K2 0905Moonshot AI95
03Claude 3.5 SonnetAnthropic94
04GPT-5OpenAI93
05Kimi K2 InstructMoonshot AI93
06Qwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team93
07o1-miniOpenAI92
08Sarvam-30BSarvam AI92
09Claude 3.5 SonnetAnthropic92
10Mistral Large 2Mistral AI92
11Qwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team92
12GPT-4oOpenAI90
13Granite 3.3 8B InstructIBM90
14Granite 3.3 8B BaseIBM90
15Gemini DiffusionGoogle90
16Nova ProAmazon89
17DeepSeek-V2.5DeepSeek89
18Llama 3.1 405B InstructMeta89
19LongCat-Flash-ChatMeituan88
20Mistral Small 3.1 24B InstructMistral AI88
21Qwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team88
22Llama 3.3 70B InstructMeta88
23Qwen2.5 32B InstructAlibaba Cloud / Qwen Team88
24Grok-2xAI88
25o1OpenAI88
26Claude 3.5 HaikuAnthropic88
27GPT-4.5OpenAI88
28Gemma 3 27BGoogle88
29GPT-4o miniOpenAI87
30GPT-4 TurboOpenAI87
31Qwen2.5 72B InstructAlibaba Cloud / Qwen Team87
32Qwen2 72B InstructAlibaba Cloud / Qwen Team86
33Grok-2 minixAI86
34Gemma 3 12BGoogle85
35Nova LiteAmazon85
36Claude 3 OpusAnthropic85
37Mistral Small 3 24B InstructMistral AI85
38Qwen2.5 7B InstructAlibaba Cloud / Qwen Team85
39Gemini 1.5 ProGoogle84
40Qwen2.5 14B InstructAlibaba Cloud / Qwen Team84
41Phi 4Microsoft83
42IBM Granite 4.0 Tiny PreviewIBM82
43Nova MicroAmazon81
44Codestral-22BMistral AI81
45Llama 3.1 70B InstructMeta81
46Qwen2 7B InstructAlibaba Cloud / Qwen Team80
47Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team79
48Claude 3 HaikuAnthropic76
49Gemma 3n E4B InstructedGoogle75
50Gemma 3n E4B Instructed LiteRT PreviewGoogle75
51Gemini 1.5 FlashGoogle74
52Grok-1.5xAI74
53Claude 3 SonnetAnthropic73
54Llama 3.1 8B InstructMeta73
55Pixtral-12BMistral AI72
56Gemma 3 4BGoogle71
57Phi-3.5-MoE-instructMicrosoft71
58GPT-3.5 TurboOpenAI68
59GPT-4OpenAI67
60Gemma 3n E2B InstructedGoogle67

60 of 66 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC