Global-MMLU-Lite Benchmark Leaderboard

Multilingual & MultimodalLive

Category

Multilingual & Multimodal

Source

Artificial Analysis

evaluation of record

Models covered

14

in our data

Data status

Live

Top score

89.2%

best on record

Top model

Gemini 2.5 Pro Preview 06-05

Google

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A lightweight, multilingual version of MMLU, designed to evaluate knowledge and reasoning skills across a diverse range of languages and cultural contexts.

Leaderboard

Top 14 of 14 models we hold a score for.

1Gemini 2.5 Pro Preview 06-05Google
89.2%2Gemini 2.5 ProGoogle
88.6%3Gemini 2.5 FlashGoogle
88.4%4Gemini 2.5 Flash-LiteGoogle
81.1%5Gemini 2.0 Flash-LiteGoogle
78.2%6Gemma 3 27BGoogle
75.1%7Gemma 3 12BGoogle
69.5%8Gemini DiffusionGoogle
69.1%9Gemma 3n E4B Instructed LiteRT PreviewGoogle
64.5%10Gemma 3n E4B InstructedGoogle
64.5%11Gemma 3n E2B Instructed LiteRTPreviewGoogle
59.0%12Gemma 3n E2B InstructedGoogle
59.0%13Gemma 3 4BGoogle
54.5%14Gemma 3 1BGoogle
34.2%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

A trimmed, translated MMLU built to check whether a model’s knowledge survives the trip out of English, across a spread of languages and cultural contexts. Higher is better, but the number worth watching is not the average — it is the spread between a model’s best and worst languages, which is where multilingual weakness actually shows up. Translated benchmarks inherit the quality of their translations, and some questions are culturally specific in ways that make the "correct" answer arguable outside the context they came from. Coverage is deliberately light, so treat it as a screening test rather than a serious multilingual evaluation.

© 2026 NYSGPT2525 LLC