Models scored
34
evaluated
Modality
text
Category
code
+1 more
Published
2025
arxiv.org
Citations
118
Semantic Scholar
Influential
12
citations
References
43
cited works
Venue
arXiv.org
published in
Abstract
Daoguang Zan, Zhirong Huang, Wei Liu, Hanwu Chen, et al. (+15)
The task of issue resolving is to modify a codebase to generate a patch that addresses a given issue. However, existing benchmarks, such as SWE-bench, focus almost exclusively on Python, making them insufficient for evaluating Large Language Models (LLMs) across diverse software ecosystems. To address this, we introduce a multilingual issue-resolving benchmark, called Multi-SWE-bench, covering Java, TypeScript, JavaScript, Go, Rust, C, and C++. It includes a total of 1,632 high-quality instances, which were carefully annotated from 2,456 candidates by 68 expert annotators, ensuring that the benchmark can provide an accurate and reliable evaluation. Based on Multi-SWE-bench, we evaluate a series of state-of-the-art models using three representative methods (Agentless, SWE-agent, and OpenHands) and present a comprehensive analysis with key empirical insights. In addition, we launch a Multi-SWE-RL open-source community, aimed at building large-scale reinforcement learning (RL) training datasets for issue-resolving tasks. As an initial contribution, we release a set of 4,723 well-structured instances spanning seven programming languages, laying a solid foundation for RL research in this domain. More importantly, we open-source our entire data production pipeline, along with detailed tutorials, encouraging the open-source community to continuously contribute and expand the dataset. We envision our Multi-SWE-bench and the ever-growing Multi-SWE-RL community as catalysts for advancing RL toward its full potential, bringing us one step closer to the dawn of AGI.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Claude Mythos Preview | Anthropic | 87 |
| 02 | Claude Opus 4.8 | Anthropic | 84 |
| 03 | Claude Sonnet 5 | Anthropic | 78 |
| 04 | Qwen3.7 Max | Alibaba Cloud / Qwen Team | 78 |
| 05 | Claude Opus 4.6 | Anthropic | 78 |
| 06 | Kimi K2.6 | Moonshot AI | 77 |
| 07 | MiniMax M2.7 | MiniMax | 77 |
| 08 | DeepSeek-V4-Pro-Max | DeepSeek | 76 |
| 09 | Qwen3.7-Plus | Alibaba Cloud / Qwen Team | 76 |
| 10 | Hy3 | Tencent | 76 |
| 11 | Qwen3.6 Plus | Alibaba Cloud / Qwen Team | 74 |
| 12 | DeepSeek-V4-Flash-Max | DeepSeek | 73 |
| 13 | Kimi K2.5 | Moonshot AI | 73 |
| 14 | MiniMax M2.1 | MiniMax | 73 |
| 15 | MiMo-V2-Pro | Xiaomi | 72 |
| 16 | MiMo-V2-Flash | Xiaomi | 72 |
| 17 | Qwen3.6-27B | Alibaba Cloud / Qwen Team | 71 |
| 18 | DeepSeek-V3.2 (Thinking) | DeepSeek | 70 |
| 19 | DeepSeek-V3.2 | DeepSeek | 70 |
| 20 | Qwen3.5-397B-A17B | Alibaba Cloud / Qwen Team | 69 |
| 21 | Nemotron 3 Ultra (550B A55B) | NVIDIA | 68 |
| 22 | Qwen3.6-35B-A3B | Alibaba Cloud / Qwen Team | 67 |
| 23 | GLM-4.7 | Zhipu AI | 67 |
| 24 | MAI-Code-1-Flash | Microsoft | 66 |
| 25 | Kimi K2-Thinking-0905 | Moonshot AI | 61 |
| 26 | DeepSeek-V3.2-Exp | DeepSeek | 58 |
| 27 | MiniMax M2 | MiniMax | 56 |
| 28 | Qwen3-Coder 480B A35B Instruct | Alibaba Cloud / Qwen Team | 55 |
| 29 | DeepSeek-V3.1 | DeepSeek | 55 |
| 30 | Kimi K2-Instruct-0905 | Moonshot AI | 47 |
| 31 | Kimi K2 Instruct | Moonshot AI | 47 |
| 32 | Nemotron 3 Super (120B A12B) | NVIDIA | 46 |
| 33 | LongCat-Flash-Lite | Meituan | 38 |
| 34 | DeepSeek-R1-0528 | DeepSeek | 31 |
34 of 34 models · score normalized 0–100 where available