Models scored
104
evaluated
Modality
text
Category
code
+2 more
Published
2023
arxiv.org
Citations
2,478
Semantic Scholar
Influential
395
citations
References
64
cited works
Venue
International Conference on Learning Representations
published in
Abstract
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, et al. (+3)
Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of $2,294$ software engineering problems drawn from real GitHub issues and corresponding pull requests across $12$ popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere $1.96$% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Claude Fable 5 | Anthropic | 95 |
| 02 | Claude Mythos Preview | Anthropic | 94 |
| 03 | Claude Opus 4.8 | Anthropic | 89 |
| 04 | Claude Opus 4.7 | Anthropic | 88 |
| 05 | Claude Sonnet 5 | Anthropic | 85 |
| 06 | Claude Opus 4.5 | Anthropic | 81 |
| 07 | Claude Opus 4.6 | Anthropic | 81 |
| 08 | Gemini 3.1 Pro | 81 | |
| 09 | DeepSeek-V4-Pro-Max | DeepSeek | 81 |
| 10 | MiniMax M3 | MiniMax | 81 |
| 11 | Qwen3.7 Max | Alibaba Cloud / Qwen Team | 80 |
| 12 | Kimi K2.6 | Moonshot AI | 80 |
| 13 | MiniMax M2.5 | MiniMax | 80 |
| 14 | GPT-5.2 | OpenAI | 80 |
| 15 | Claude Sonnet 4.6 | Anthropic | 80 |
| 16 | DeepSeek-V4-Flash-Max | DeepSeek | 79 |
| 17 | MiMo-V2.5-Pro | Xiaomi | 79 |
| 18 | Qwen3.6 Plus | Alibaba Cloud / Qwen Team | 79 |
| 19 | Gemini 3 Flash | 78 | |
| 20 | Hy3 | Tencent | 78 |
| 21 | MiMo-V2-Pro | Xiaomi | 78 |
| 22 | GLM-5 | Zhipu AI | 78 |
| 23 | Qwen3.7-Plus | Alibaba Cloud / Qwen Team | 78 |
| 24 | Mistral Medium 3.5 | Mistral AI | 78 |
| 25 | Muse Spark | Meta | 77 |
| 26 | Qwen3.6-27B | Alibaba Cloud / Qwen Team | 77 |
| 27 | Kimi K2.5 | Moonshot AI | 77 |
| 28 | Seed 2.0 Pro | ByteDance | 77 |
| 29 | Qwen3.5-397B-A17B | Alibaba Cloud / Qwen Team | 76 |
| 30 | GPT-5.1 | OpenAI | 76 |
| 31 | GPT-5.1 Instant | OpenAI | 76 |
| 32 | GPT-5.1 Thinking | OpenAI | 76 |
| 33 | Gemini 3 Pro | 76 | |
| 34 | GPT-5 | OpenAI | 75 |
| 35 | MiMo-V2-Omni | Xiaomi | 75 |
| 36 | Claude Opus 4.1 | Anthropic | 75 |
| 37 | GPT-5 Codex | OpenAI | 75 |
| 38 | Step-3.5-Flash | StepFun | 74 |
| 39 | GLM-4.7 | Zhipu AI | 74 |
| 40 | GPT-5.1 Codex | OpenAI | 74 |
| 41 | Seed 2.0 Lite | ByteDance | 74 |
| 42 | MAI-Thinking-1 | Microsoft | 74 |
| 43 | Qwen3.6-35B-A3B | Alibaba Cloud / Qwen Team | 73 |
| 44 | MiMo-V2-Flash | Xiaomi | 73 |
| 45 | Claude Haiku 4.5 | Anthropic | 73 |
| 46 | DeepSeek-V3.2 | DeepSeek | 73 |
| 47 | DeepSeek-V3.2 (Thinking) | DeepSeek | 73 |
| 48 | DeepSeek-V3.2-Speciale | DeepSeek | 73 |
| 49 | Claude Sonnet 4 | Anthropic | 73 |
| 50 | Claude Opus 4 | Anthropic | 73 |
| 51 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 72 |
| 52 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 72 |
| 53 | MAI-Code-1-Flash | Microsoft | 72 |
| 54 | Kimi K2-Thinking-0905 | Moonshot AI | 71 |
| 55 | Grok Code Fast 1 | xAI | 71 |
| 56 | Nemotron 3 Ultra (550B A55B) | NVIDIA | 71 |
| 57 | Claude 3.7 Sonnet | Anthropic | 70 |
| 58 | LongCat-Flash-Thinking-2601 | Meituan | 70 |
| 59 | Nova 2 Pro | Amazon | 70 |
| 60 | Qwen3-Coder 480B A35B Instruct | Alibaba Cloud / Qwen Team | 70 |
60 of 104 models · score normalized 0–100 where available