Exploring Text-to-Text Transformers for English to Hinglish Machine Translation with Synthetic Code-Mixing

We describe models focused at the understudied problem of translating between\nmonolingual and code-mixed language pairs. More specifically, we offer a wide\nrange of models that convert monolingual English text into Hinglish (code-mixed\nHindi and English). Given the recent success of pretrained language models, we\nalso test the utility of two recent Transformer-based encoder-decoder models\n(i.e., mT5 and mBART) on the task finding both to work well. Given the paucity\nof training data for code-mixing, we also propose a dependency-free method for\ngenerating code-mixed texts from bilingual distributed representations that we\nexploit for improving language model performance. In particular, armed with\nthis additional data, we adopt a curriculum learning approach where we first\nfinetune the language models on synthetic data then on gold code-mixed data. We\nfind that, although simple, our synthetic code-mixing method is competitive\nwith (and in some cases is even superior to) several standard methods\n(backtranslation, method based on equivalence constraint theory) under a\ndiverse set of conditions. Our work shows that the mT5 model, finetuned\nfollowing the curriculum learning procedure, achieves best translation\nperformance (12.67 BLEU). Our models place first in the overall ranking of the\nEnglish-Hinglish official shared task.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC