A Machine Learning-Based Evaluation of English-to-Sinhala Translation: Comparing Google Translate, Large Language Models, and Human Translators
This study presents a two-stage evaluation framework for English-to-Sinhala translation across seven systems: DeepSeek-v1, Gemini-2.5, Google Translate, GPT-4.1, GPT-5-mini, Grok-fast-4.1, and Claude Sonnet-4.5. In the first stage, source-side linguistic features are used to predict human-rated translation quality on a 0–10 scale using supervised regression models. The feature set combines TF-IDF word and character n-grams with handcrafted sentence-level linguistic features. In the second stage, system outputs are compared against human reference translations using BLEU, METEOR, chrF, TER, and character-level precision, recall, and F1. Experiments were conducted on a dataset of 2,308 English sentences, each translated by seven systems and evaluated by five bilingual human assessors. To strengthen reproducibility, this study reports inter-rater agreement, prompt settings, output coverage, and verification checks for automatic metrics. The results suggest that Gemini-2.5 performs strongly under the present evaluation setup, although the final ranking is interpreted cautiously because automatic metrics for English-to-Sinhala translation may be sensitive to tokenization, Unicode normalization, reference wording, and output formatting. The quality estimation models achieve mean absolute errors between 0.7240 and 0.9646, indicating that source-side features provide a useful, though incomplete, signal of translation difficulty. The findings provide an empirical comparison of modern translation systems for a low-resource language pair and propose a reproducible framework for translation assessment in morphologically rich languages.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex