MGSM

Language Models are Multilingual Chain-of-Thought Reasoners

Models scored

31

evaluated

Modality

text

Category

math

+1 more

Published

2022

arxiv.org

Citations

591

Semantic Scholar

Influential

112

citations

References

53

cited works

Venue

International Conference on Learning Representations

published in

Abstract

Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, et al. (+8)

We evaluate the reasoning abilities of large language models in multilingual settings. We introduce the Multilingual Grade School Math (MGSM) benchmark, by manually translating 250 grade-school math problems from the GSM8K dataset (Cobbe et al., 2021) into ten typologically diverse languages. We find that the ability to solve MGSM problems via chain-of-thought prompting emerges with increasing model scale, and that models have strikingly strong multilingual reasoning abilities, even in underrepresented languages such as Bengali and Swahili. Finally, we show that the multilingual reasoning abilities of language models extend to other tasks such as commonsense reasoning and word-in-context semantic judgment. The MGSM benchmark is publicly available at https://github.com/google-research/url-nlp.

Search

#ModelLabScore
01Llama 4 MaverickMeta92
02o3-miniOpenAI92
03Claude 3.5 SonnetAnthropic92
04Claude 3.5 SonnetAnthropic92
05Llama 3.3 70B InstructMeta91
06o1-previewOpenAI91
07Claude 3 OpusAnthropic91
08Llama 4 ScoutMeta91
09GPT-4oOpenAI91
10o1OpenAI89
11GPT-4 TurboOpenAI89
12Gemini 1.5 ProGoogle88
13GPT-4o miniOpenAI87
14Llama 3.2 90B InstructMeta87
15Claude 3.5 HaikuAnthropic86
16Qwen3 235B A22BAlibaba Cloud / Qwen Team84
17Claude 3 SonnetAnthropic84
18Gemini 1.5 FlashGoogle83
19Phi 4Microsoft81
20Claude 3 HaikuAnthropic75
21GPT-4OpenAI75
22Llama 3.2 11B InstructMeta69
23Gemma 3n E4B InstructedGoogle67
24Phi 4 MiniMicrosoft64
25Gemma 3n E4B Instructed LiteRT PreviewGoogle61
26Phi-3.5-MoE-instructMicrosoft59
27Llama 3.2 3B InstructMeta58
28GPT-3.5 TurboOpenAI56
29Gemma 3n E2B InstructedGoogle53
30Gemma 3n E2B Instructed LiteRT (Preview)Google53
31Phi-3.5-mini-instructMicrosoft48

31 of 31 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC