SWE-Bench Verified

Can Language Models Resolve Real-World GitHub Issues?

Models scored

104

evaluated

Modality

text

Category

code

+2 more

Published

2023

arxiv.org

Citations

2,478

Semantic Scholar

Influential

395

citations

References

64

cited works

Venue

International Conference on Learning Representations

published in

Abstract

Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, et al. (+3)

Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of $2,294$ software engineering problems drawn from real GitHub issues and corresponding pull requests across $12$ popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere $1.96$% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.

codefrontend developmentreasoning

Search

#ModelLabScore
01Claude Fable 5Anthropic95
02Claude Mythos PreviewAnthropic94
03Claude Opus 4.8Anthropic89
04Claude Opus 4.7Anthropic88
05Claude Sonnet 5Anthropic85
06Claude Opus 4.5Anthropic81
07Claude Opus 4.6Anthropic81
08Gemini 3.1 ProGoogle81
09DeepSeek-V4-Pro-MaxDeepSeek81
10MiniMax M3MiniMax81
11Qwen3.7 MaxAlibaba Cloud / Qwen Team80
12Kimi K2.6Moonshot AI80
13MiniMax M2.5MiniMax80
14GPT-5.2OpenAI80
15Claude Sonnet 4.6Anthropic80
16DeepSeek-V4-Flash-MaxDeepSeek79
17MiMo-V2.5-ProXiaomi79
18Qwen3.6 PlusAlibaba Cloud / Qwen Team79
19Gemini 3 FlashGoogle78
20Hy3Tencent78
21MiMo-V2-ProXiaomi78
22GLM-5Zhipu AI78
23Qwen3.7-PlusAlibaba Cloud / Qwen Team78
24Mistral Medium 3.5Mistral AI78
25Muse SparkMeta77
26Qwen3.6-27BAlibaba Cloud / Qwen Team77
27Kimi K2.5Moonshot AI77
28Seed 2.0 ProByteDance77
29Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team76
30GPT-5.1OpenAI76
31GPT-5.1 InstantOpenAI76
32GPT-5.1 ThinkingOpenAI76
33Gemini 3 ProGoogle76
34GPT-5OpenAI75
35MiMo-V2-OmniXiaomi75
36Claude Opus 4.1Anthropic75
37GPT-5 CodexOpenAI75
38Step-3.5-FlashStepFun74
39GLM-4.7Zhipu AI74
40GPT-5.1 CodexOpenAI74
41Seed 2.0 LiteByteDance74
42MAI-Thinking-1Microsoft74
43Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team73
44MiMo-V2-FlashXiaomi73
45Claude Haiku 4.5Anthropic73
46DeepSeek-V3.2DeepSeek73
47DeepSeek-V3.2 (Thinking)DeepSeek73
48DeepSeek-V3.2-SpecialeDeepSeek73
49Claude Sonnet 4Anthropic73
50Claude Opus 4Anthropic73
51Qwen3.5-27BAlibaba Cloud / Qwen Team72
52Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team72
53MAI-Code-1-FlashMicrosoft72
54Kimi K2-Thinking-0905Moonshot AI71
55Grok Code Fast 1xAI71
56Nemotron 3 Ultra (550B A55B)NVIDIA71
57Claude 3.7 SonnetAnthropic70
58LongCat-Flash-Thinking-2601Meituan70
59Nova 2 ProAmazon70
60Qwen3-Coder 480B A35B InstructAlibaba Cloud / Qwen Team70

60 of 104 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC