Maastricht University at AMIYA: Adapting LLMs for Dialectal Arabic using Fine-tuning and MBR Decoding

Large Language Models (LLMs) are becoming increasingly multilingual, supporting hundreds of languages, especially high resource ones. Unfortunately, Dialect variations are still underrepresented due to limited data and linguistic variation. In this work, we adapt a pre-trained LLM to improve dialectal performance. Specifically, we use Low Rank Adaptation (LoRA) fine-tuning on monolingual and English Dialect parallel data, adapter merging and dialect-aware MBR decoding to improve dialectal fidelity generation and translation. Experiments on Syrian, Moroccan, and Saudi Arabic show that merging and MBR improve dialectal fidelity while preserving semantic accuracy. This combination provides a compact and effective framework for robust dialectal Arabic generation.

Paper

References (19)

11Saudial: Saudi arabic di-alects game localization dataset2025 · . Parallel Saudi dialect text dataset (English, MSA, Saudi varieties) with cultural context, age ratings, and dialect notes
12Inception, cerebras and mbzuai release Jais 2 – the next generation arabic open-weight llm2025 · mbzuai

Scroll for more · 7 remaining

Similar papers

© 2026 NYSGPT2525 LLC