Benchmarking Next-Generation Frontier Models for Document-Level Arabic-Thai Medical Translation: A Reliability Study of LLM-as-a-Judge

This research targets Arabic-to-Thai medical translation, a low-resource and linguistically distant pair critical for public health. We utilize the OpenWHO dataset, a resource comprising 293 documents. This dataset is shielded from webcrawling to minimize training data contamination and ensure rigorous zero-shot evaluation. We conduct a comparative analysis between 2026-era frontier models including GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5 and industry-standard NMT represented by Google Translate. Using an LLM-as-a-Judge framework with an AI jury of efficient reasoning models such as GPT-5.1, Gemini 3 Flash, and Claude Sonnet 4.5, we performed 3,516 evaluations of fidelity, fluency, and cultural appropriateness. Results demonstrate that frontier models outperform traditional NMT across all dimensions. While high exact agreement suggests LLM-as-a-Judge frameworks are promising for scalable evaluation, human validation reveals persistent opportunities for improving AI-human alignment, necessitating targeted oversight for safety-critical applications.

Paper

Full text

PDF

Benchmarking Next-Generation Frontier Models for Document-Level Arabic-Thai Medical Translation: A Reliability Study of LLM-as-a-Judge

Semantic Scholar · 2026

Abstract

This research targets Arabic-to-Thai medical translation, a low-resource and linguistically distant pair critical for public health. We utilize the OpenWHO dataset, a resource comprising 293 documents. This dataset is shielded from webcrawling to minimize training data contamination and ensure rigorous zero-shot evaluation. We conduct a comparative analysis between 2026-era frontier models including GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5 and industry-standard NMT represented by Google Translate. Using an LLM-as-a-Judge framework with an AI jury of efficient reasoning models such as GPT-5.1, Gemini 3 Flash, and Claude Sonnet 4.5, we performed 3,516 evaluations of fidelity, fluency, and cultural appropriateness. Results demonstrate that frontier models outperform traditional NMT across all dimensions. While high exact agreement suggests LLM-as-a-Judge frameworks are promising for scalable evaluation, human validation reveals persistent opportunities for improving AI-human alignment, necessitating targeted oversight for safety-critical applications.

Similar papers

© 2026 NYSGPT2525 LLC