Models scored
31
evaluated
Modality
text
Category
math
+1 more
Published
2022
arxiv.org
Citations
591
Semantic Scholar
Influential
112
citations
References
53
cited works
Venue
International Conference on Learning Representations
published in
Abstract
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, et al. (+8)
We evaluate the reasoning abilities of large language models in multilingual settings. We introduce the Multilingual Grade School Math (MGSM) benchmark, by manually translating 250 grade-school math problems from the GSM8K dataset (Cobbe et al., 2021) into ten typologically diverse languages. We find that the ability to solve MGSM problems via chain-of-thought prompting emerges with increasing model scale, and that models have strikingly strong multilingual reasoning abilities, even in underrepresented languages such as Bengali and Swahili. Finally, we show that the multilingual reasoning abilities of language models extend to other tasks such as commonsense reasoning and word-in-context semantic judgment. The MGSM benchmark is publicly available at https://github.com/google-research/url-nlp.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Llama 4 Maverick | Meta | 92 |
| 02 | o3-mini | OpenAI | 92 |
| 03 | Claude 3.5 Sonnet | Anthropic | 92 |
| 04 | Claude 3.5 Sonnet | Anthropic | 92 |
| 05 | Llama 3.3 70B Instruct | Meta | 91 |
| 06 | o1-preview | OpenAI | 91 |
| 07 | Claude 3 Opus | Anthropic | 91 |
| 08 | Llama 4 Scout | Meta | 91 |
| 09 | GPT-4o | OpenAI | 91 |
| 10 | o1 | OpenAI | 89 |
| 11 | GPT-4 Turbo | OpenAI | 89 |
| 12 | Gemini 1.5 Pro | 88 | |
| 13 | GPT-4o mini | OpenAI | 87 |
| 14 | Llama 3.2 90B Instruct | Meta | 87 |
| 15 | Claude 3.5 Haiku | Anthropic | 86 |
| 16 | Qwen3 235B A22B | Alibaba Cloud / Qwen Team | 84 |
| 17 | Claude 3 Sonnet | Anthropic | 84 |
| 18 | Gemini 1.5 Flash | 83 | |
| 19 | Phi 4 | Microsoft | 81 |
| 20 | Claude 3 Haiku | Anthropic | 75 |
| 21 | GPT-4 | OpenAI | 75 |
| 22 | Llama 3.2 11B Instruct | Meta | 69 |
| 23 | Gemma 3n E4B Instructed | 67 | |
| 24 | Phi 4 Mini | Microsoft | 64 |
| 25 | Gemma 3n E4B Instructed LiteRT Preview | 61 | |
| 26 | Phi-3.5-MoE-instruct | Microsoft | 59 |
| 27 | Llama 3.2 3B Instruct | Meta | 58 |
| 28 | GPT-3.5 Turbo | OpenAI | 56 |
| 29 | Gemma 3n E2B Instructed | 53 | |
| 30 | Gemma 3n E2B Instructed LiteRT (Preview) | 53 | |
| 31 | Phi-3.5-mini-instruct | Microsoft | 48 |
31 of 31 models · score normalized 0–100 where available