Multilingual transfer of acoustic word embeddings improves when training on languages related to the target zero-resource language
Acoustic word embedding models map variable duration speech segments to fixed\ndimensional vectors, enabling efficient speech search and discovery. Previous\nwork explored how embeddings can be obtained in zero-resource settings where no\nlabelled data is available in the target language. The current best approach\nuses transfer learning: a single supervised multilingual model is trained using\nlabelled data from multiple well-resourced languages and then applied to a\ntarget zero-resource language (without fine-tuning). However, it is still\nunclear how the specific choice of training languages affect downstream\nperformance. Concretely, here we ask whether it is beneficial to use training\nlanguages related to the target. Using data from eleven languages spoken in\nSouthern Africa, we experiment with adding data from different language\nfamilies while controlling for the amount of data per language. In word\ndiscrimination and query-by-example search evaluations, we show that training\non languages from the same family gives large improvements. Through\nfiner-grained analysis, we show that training on even just a single related\nlanguage gives the largest gain. We also find that adding data from unrelated\nlanguages generally doesn't hurt performance.\n
Paper
References (44)
Scroll for more · 32 remaining