Machine translation for low-resource languages faces a scientific problem of cross-language semantic paraphrasing in a lack of structured language resource. The exploration of this problem has important theoretical and application value and is also a challenging research hotspot at present. We address the specific low-resource machine translation issue from Indonesian to Chinese, propose a language resource extension method based on cognate parallel corpus, and train a modified neural machine translation (NMT) model by mixing parallel corpus from cognate language. This modified model achieved 20.30 BLEU4 score in the experiment of Indonesian-Chinese machine translation. The manual analysis after simple random sampling of experimental results finds that the effect of the modified NMT is comparable to that of the current Google translation. The experimental results prove that the cognate parallel corpus can improve the low-resource language NMT effectively, which mainly depends on the morphological similarity and semantic equivalence between the cognate languages.
Paper
Full text
Language Resource Extension for Indonesian-Chinese Machine Translation
Semantic Scholar · Computer Science · 2018
Abstract
Machine translation for low-resource languages faces a scientific problem of cross-language semantic paraphrasing in a lack of structured language resource. The exploration of this problem has important theoretical and application value and is also a challenging research hotspot at present. We address the specific low-resource machine translation issue from Indonesian to Chinese, propose a language resource extension method based on cognate parallel corpus, and train a modified neural machine translation (NMT) model by mixing parallel corpus from cognate language. This modified model achieved 20.30 BLEU4 score in the experiment of Indonesian-Chinese machine translation. The manual analysis after simple random sampling of experimental results finds that the effect of the modified NMT is comparable to that of the current Google translation. The experimental results prove that the cognate parallel corpus can improve the low-resource language NMT effectively, which mainly depends on the morphological similarity and semantic equivalence between the cognate languages.