MBPP

Program Synthesis with Large Language Models

Models scored

33

evaluated

Modality

text

Category

general

+1 more

Published

2021

arxiv.org

Citations

3,804

Semantic Scholar

Influential

580

citations

References

106

cited works

Venue

arXiv.org

published in

Abstract

Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, et al. (+7)

This paper explores the limits of the current generation of large language models for program synthesis in general purpose programming languages. We evaluate a collection of such models (with between 244M and 137B parameters) on two new benchmarks, MBPP and MathQA-Python, in both the few-shot and fine-tuning regimes. Our benchmarks are designed to measure the ability of these models to synthesize short Python programs from natural language descriptions. The Mostly Basic Programming Problems (MBPP) dataset contains 974 programming tasks, designed to be solvable by entry-level programmers. The MathQA-Python dataset, a Python version of the MathQA benchmark, contains 23914 problems that evaluate the ability of the models to synthesize code from more complex text. On both datasets, we find that synthesis performance scales log-linearly with model size. Our largest models, even without finetuning on a code dataset, can synthesize solutions to 59.6 percent of the problems from MBPP using few-shot learning with a well-designed prompt. Fine-tuning on a held-out portion of the dataset improves performance by about 10 percentage points across most model sizes. On the MathQA-Python dataset, the largest fine-tuned model achieves 83.8 percent accuracy. Going further, we study the model's ability to engage in dialog about code, incorporating human feedback to improve its solutions. We find that natural language feedback from a human halves the error rate compared to the model's initial prediction. Additionally, we conduct an error analysis to shed light on where these models fall short and what types of programs are most difficult to generate. Finally, we explore the semantic grounding of these models by fine-tuning them to predict the results of program execution. We find that even our best models are generally unable to predict the output of a program given a specific input.

generalreasoning

Search

#ModelLabScore
01Sarvam-30BSarvam AI93
02Llama-3.3 Nemotron Super 49B v1NVIDIA91
03Qwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team90
04MiniCPM-SALAOpenBMB89
05Qwen2.5 72B InstructAlibaba Cloud / Qwen Team88
06Llama 3.1 Nemotron Nano 8B V1NVIDIA85
07Qwen2.5 32B InstructAlibaba Cloud / Qwen Team84
08Qwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team84
09Qwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team84
10Qwen2.5 14B InstructAlibaba Cloud / Qwen Team82
11Qwen3 235B A22BAlibaba Cloud / Qwen Team81
12Phi-3.5-MoE-instructMicrosoft81
13Qwen2 72B InstructAlibaba Cloud / Qwen Team80
14Qwen2.5 7B InstructAlibaba Cloud / Qwen Team79
15Codestral-22BMistral AI78
16Llama 4 MaverickMeta78
17Gemini DiffusionGoogle76
18Mistral Small 3.1 24B InstructMistral AI75
19Gemma 3 27BGoogle74
20Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team73
21Gemma 3 12BGoogle73
22Mistral Small 3 24B BaseMistral AI70
23Phi-3.5-mini-instructMicrosoft70
24Llama 4 ScoutMeta68
25Qwen2 7B InstructAlibaba Cloud / Qwen Team67
26Gemma 3n E4B InstructedGoogle64
27Gemma 3n E4B Instructed LiteRT PreviewGoogle64
28Gemma 3 4BGoogle63
29Gemma 2 27BGoogle63
30Gemma 3n E2B Instructed LiteRT (Preview)Google57
31Gemma 3n E2B InstructedGoogle57
32Gemma 2 9BGoogle52
33Gemma 3 1BGoogle35

33 of 33 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC