Investigating Language Impact in Bilingual Approaches for Computational Language Documentation

For endangered languages, data collection campaigns have to accommodate the\nchallenge that many of them are from oral tradition, and producing\ntranscriptions is costly. Therefore, it is fundamental to translate them into a\nwidely spoken language to ensure interpretability of the recordings. In this\npaper we investigate how the choice of translation language affects the\nposterior documentation work and potential automatic approaches which will work\non top of the produced bilingual corpus. For answering this question, we use\nthe MaSS multilingual speech corpus (Boito et al., 2020) for creating 56\nbilingual pairs that we apply to the task of low-resource unsupervised word\nsegmentation and alignment. Our results highlight that the choice of language\nfor translation influences the word segmentation performance, and that\ndifferent lexicons are learned by using different aligned translations. Lastly,\nthis paper proposes a hybrid approach for bilingual word segmentation,\ncombining boundary clues extracted from a non-parametric Bayesian model\n(Goldwater et al., 2009a) with the attentional word segmentation neural model\nfrom Godard et al. (2018). Our results suggest that incorporating these clues\ninto the neural models' input representation increases their translation and\nalignment quality, specially for challenging language pairs.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC