The Multilingual Evaluation Paradox: Ramoju Potential Across Languages, Scripts, and 2.37 Billion Speakers

Technical Report XARL-2026-06. We evaluate Llama-3-8B-Instruct on the Multilingual GSM8K (MGSM) benchmark across seven languages covering 2.37 billion native speakers and find that every non-English language exhibits higher Ramoju Potential than English without exception. We introduce the Multilingual Suppression Index (MSI), finding Chinese reaches MSI=2.87x — meaning 1.4 billion Chinese speakers face an evaluation gap nearly three times larger than English speakers. We identify three instruction-response patterns: instruction-hurt (English), instruction-positive (Chinese, Japanese, Bengali, Swahili, Telugu), and instruction-amplifying (Russian). We establish the Multilingual Evaluation Paradox: format suppression scales inversely with language resource richness, compounding linguistic inequality in AI evaluation. Sixth in the XARL evaluation series.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC