Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine Translation
Despite being the seventh most widely spoken language in the world, Bengali\nhas received much less attention in machine translation literature due to being\nlow in resources. Most publicly available parallel corpora for Bengali are not\nlarge enough; and have rather poor quality, mostly because of incorrect\nsentence alignments resulting from erroneous sentence segmentation, and also\nbecause of a high volume of noise present in them. In this work, we build a\ncustomized sentence segmenter for Bengali and propose two novel methods for\nparallel corpus creation on low-resource setups: aligner ensembling and batch\nfiltering. With the segmenter and the two methods combined, we compile a\nhigh-quality Bengali-English parallel corpus comprising of 2.75 million\nsentence pairs, more than 2 million of which were not available before.\nTraining on neural models, we achieve an improvement of more than 9 BLEU score\nover previous approaches to Bengali-English machine translation. We also\nevaluate on a new test set of 1000 pairs made with extensive quality control.\nWe release the segmenter, parallel corpus, and the evaluation set, thus\nelevating Bengali from its low-resource status. To the best of our knowledge,\nthis is the first ever large scale study on Bengali-English machine\ntranslation. We believe our study will pave the way for future research on\nBengali-English machine translation as well as other low-resource languages.\nOur data and code are available at https://github.com/csebuetnlp/banglanmt.\n
Paper
References (57)
Scroll for more · 38 remaining