Crowdsourcing Parallel Corpus for English-Oromo Neural Machine Translation using Community Engagement Platform

Even though Afaan Oromo is the most widely spoken language in the Cushitic\nfamily by more than fifty million people in the Horn and East Africa, it is\nsurprisingly resource-scarce from a technological point of view. The increasing\namount of various useful documents written in English language brings to\ninvestigate the machine that can translate those documents and make it easily\naccessible for local language. The paper deals with implementing a translation\nof English to Afaan Oromo and vice versa using Neural Machine Translation. But\nthe implementation is not very well explored due to the limited amount and\ndiversity of the corpus. However, using a bilingual corpus of just over 40k\nsentence pairs we have collected, this study showed a promising result. About a\nquarter of this corpus is collected via Community Engagement Platform (CEP)\nthat was implemented to enrich the parallel corpus through crowdsourcing\ntranslations.\n

Paper

References (15)

10Hiikaa - English to Afaan Oromoo Translation2020 · Retrieved from https://hika.lingogrid.com.
12EnglishAfaan Oromo Statistical Machine Translation. International Journal of Computational Linguistic (IJCL)2018

Scroll for more · 3 remaining

Similar papers

© 2026 NYSGPT2525 LLC