Category
Multilingual & Multimodal
Source
Artificial Analysis
evaluation of record
Models covered
14
in our data
Data status
Live
Top score
89.2%
best on record
Top model
Gemini 2.5 Pro Preview 06-05
Updated
2026-07-30
last ingest
A lightweight, multilingual version of MMLU, designed to evaluate knowledge and reasoning skills across a diverse range of languages and cultural contexts.
Leaderboard
Top 14 of 14 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
A trimmed, translated MMLU built to check whether a model’s knowledge survives the trip out of English, across a spread of languages and cultural contexts. Higher is better, but the number worth watching is not the average — it is the spread between a model’s best and worst languages, which is where multilingual weakness actually shows up. Translated benchmarks inherit the quality of their translations, and some questions are culturally specific in ways that make the "correct" answer arguable outside the context they came from. Coverage is deliberately light, so treat it as a screening test rather than a serious multilingual evaluation.