Central Kurdish machine translation: First large scale parallel corpus and experiments

While the computational processing of Kurdish has experienced a relative\nincrease, the machine translation of this language seems to be lacking a\nconsiderable body of scientific work. This is in part due to the lack of\nresources especially curated for this task. In this paper, we present the first\nlarge scale parallel corpus of Central Kurdish-English, Awta, containing\n229,222 pairs of manually aligned translations. Our corpus is collected from\ndifferent text genres and domains in an attempt to build more robust and\nreal-world applications of machine translation. We make a portion of this\ncorpus publicly available in order to foster research in this area. Further, we\nbuild several neural machine translation models in order to benchmark the task\nof Kurdish machine translation. Additionally, we perform extensive experimental\nanalysis of results in order to identify the major challenges that Central\nKurdish machine translation faces. These challenges include language-dependent\nand-independent ones as categorized in this paper, the first group of which are\naware of Central Kurdish linguistic properties on different morphological,\nsyntactic and semantic levels. Our best performing systems achieve 22.72 and\n16.81 in BLEU score for Ku$\\rightarrow$EN and En$\\rightarrow$Ku, respectively.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC