Crowdsourcing Parallel Corpus for English-Oromo Neural Machine Translation using Community Engagement Platform
Even though Afaan Oromo is the most widely spoken language in the Cushitic\nfamily by more than fifty million people in the Horn and East Africa, it is\nsurprisingly resource-scarce from a technological point of view. The increasing\namount of various useful documents written in English language brings to\ninvestigate the machine that can translate those documents and make it easily\naccessible for local language. The paper deals with implementing a translation\nof English to Afaan Oromo and vice versa using Neural Machine Translation. But\nthe implementation is not very well explored due to the limited amount and\ndiversity of the corpus. However, using a bilingual corpus of just over 40k\nsentence pairs we have collected, this study showed a promising result. About a\nquarter of this corpus is collected via Community Engagement Platform (CEP)\nthat was implemented to enrich the parallel corpus through crowdsourcing\ntranslations.\n
Paper
References (15)
Scroll for more · 3 remaining