SWE-bench Multilingual

A Multilingual Benchmark for Issue Resolving

Models scored

34

evaluated

Modality

text

Category

code

+1 more

Published

2025

arxiv.org

Citations

118

Semantic Scholar

Influential

12

citations

References

43

cited works

Venue

arXiv.org

published in

Abstract

Daoguang Zan, Zhirong Huang, Wei Liu, Hanwu Chen, et al. (+15)

The task of issue resolving is to modify a codebase to generate a patch that addresses a given issue. However, existing benchmarks, such as SWE-bench, focus almost exclusively on Python, making them insufficient for evaluating Large Language Models (LLMs) across diverse software ecosystems. To address this, we introduce a multilingual issue-resolving benchmark, called Multi-SWE-bench, covering Java, TypeScript, JavaScript, Go, Rust, C, and C++. It includes a total of 1,632 high-quality instances, which were carefully annotated from 2,456 candidates by 68 expert annotators, ensuring that the benchmark can provide an accurate and reliable evaluation. Based on Multi-SWE-bench, we evaluate a series of state-of-the-art models using three representative methods (Agentless, SWE-agent, and OpenHands) and present a comprehensive analysis with key empirical insights. In addition, we launch a Multi-SWE-RL open-source community, aimed at building large-scale reinforcement learning (RL) training datasets for issue-resolving tasks. As an initial contribution, we release a set of 4,723 well-structured instances spanning seven programming languages, laying a solid foundation for RL research in this domain. More importantly, we open-source our entire data production pipeline, along with detailed tutorials, encouraging the open-source community to continuously contribute and expand the dataset. We envision our Multi-SWE-bench and the ever-growing Multi-SWE-RL community as catalysts for advancing RL toward its full potential, bringing us one step closer to the dawn of AGI.

Search

#ModelLabScore
01Claude Mythos PreviewAnthropic87
02Claude Opus 4.8Anthropic84
03Claude Sonnet 5Anthropic78
04Qwen3.7 MaxAlibaba Cloud / Qwen Team78
05Claude Opus 4.6Anthropic78
06Kimi K2.6Moonshot AI77
07MiniMax M2.7MiniMax77
08DeepSeek-V4-Pro-MaxDeepSeek76
09Qwen3.7-PlusAlibaba Cloud / Qwen Team76
10Hy3Tencent76
11Qwen3.6 PlusAlibaba Cloud / Qwen Team74
12DeepSeek-V4-Flash-MaxDeepSeek73
13Kimi K2.5Moonshot AI73
14MiniMax M2.1MiniMax73
15MiMo-V2-ProXiaomi72
16MiMo-V2-FlashXiaomi72
17Qwen3.6-27BAlibaba Cloud / Qwen Team71
18DeepSeek-V3.2 (Thinking)DeepSeek70
19DeepSeek-V3.2DeepSeek70
20Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team69
21Nemotron 3 Ultra (550B A55B)NVIDIA68
22Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team67
23GLM-4.7Zhipu AI67
24MAI-Code-1-FlashMicrosoft66
25Kimi K2-Thinking-0905Moonshot AI61
26DeepSeek-V3.2-ExpDeepSeek58
27MiniMax M2MiniMax56
28Qwen3-Coder 480B A35B InstructAlibaba Cloud / Qwen Team55
29DeepSeek-V3.1DeepSeek55
30Kimi K2-Instruct-0905Moonshot AI47
31Kimi K2 InstructMoonshot AI47
32Nemotron 3 Super (120B A12B)NVIDIA46
33LongCat-Flash-LiteMeituan38
34DeepSeek-R1-0528DeepSeek31

34 of 34 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC