Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages Study

Recent research in multilingual language models (LM) has demonstrated their\nability to effectively handle multiple languages in a single model. This holds\npromise for low web-resource languages (LRL) as multilingual models can enable\ntransfer of supervision from high resource languages to LRLs. However,\nincorporating a new language in an LM still remains a challenge, particularly\nfor languages with limited corpora and in unseen scripts. In this paper we\nargue that relatedness among languages in a language family may be exploited to\novercome some of the corpora limitations of LRLs, and propose RelateLM. We\nfocus on Indian languages, and exploit relatedness along two dimensions: (1)\nscript (since many Indic scripts originated from the Brahmic script), and (2)\nsentence structure. RelateLM uses transliteration to convert the unseen script\nof limited LRL text into the script of a Related Prominent Language (RPL)\n(Hindi in our case). While exploiting similar sentence structures, RelateLM\nutilizes readily available bilingual dictionaries to pseudo translate RPL text\ninto LRL corpora. Experiments on multiple real-world benchmark datasets provide\nvalidation to our hypothesis that using a related language as pivot, along with\ntransliteration and pseudo translation based data augmentation, can be an\neffective way to adapt LMs for LRLs, rather than direct training or pivoting\nthrough English.\n

Paper

References (38)

Scroll for more · 26 remaining

Similar papers

© 2026 NYSGPT2525 LLC