Role of Language Relatedness in Multilingual Fine-tuning of Language Models: A Case Study in Indo-Aryan Languages

We explore the impact of leveraging the relatedness of languages that belong\nto the same family in NLP models using multilingual fine-tuning. We hypothesize\nand validate that multilingual fine-tuning of pre-trained language models can\nyield better performance on downstream NLP applications, compared to models\nfine-tuned on individual languages. A first of its kind detailed study is\npresented to track performance change as languages are added to a base language\nin a graded and greedy (in the sense of best boost of performance) manner;\nwhich reveals that careful selection of subset of related languages can\nsignificantly improve performance than utilizing all related languages. The\nIndo-Aryan (IA) language family is chosen for the study, the exact languages\nbeing Bengali, Gujarati, Hindi, Marathi, Oriya, Punjabi and Urdu. The script\nbarrier is crossed by simple rule-based transliteration of the text of all\nlanguages to Devanagari. Experiments are performed on mBERT, IndicBERT, MuRIL\nand two RoBERTa-based LMs, the last two being pre-trained by us. Low resource\nlanguages, such as Oriya and Punjabi, are found to be the largest beneficiaries\nof multilingual fine-tuning. Textual Entailment, Entity Classification, Section\nTitle Prediction, tasks of IndicGLUE and POS tagging form our test bed.\nCompared to monolingual fine tuning we get relative performance improvement of\nup to 150% in the downstream tasks. The surprise take-away is that for any\nlanguage there is a particular combination of other languages which yields the\nbest performance, and any additional language is in fact detrimental.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC