Improved acoustic word embeddings for zero-resource languages using multilingual transfer

Acoustic word embeddings are fixed-dimensional representations of\nvariable-length speech segments. Such embeddings can form the basis for speech\nsearch, indexing and discovery systems when conventional speech recognition is\nnot possible. In zero-resource settings where unlabelled speech is the only\navailable resource, we need a method that gives robust embeddings on an\narbitrary language. Here we explore multilingual transfer: we train a single\nsupervised embedding model on labelled data from multiple well-resourced\nlanguages and then apply it to unseen zero-resource languages. We consider\nthree multilingual recurrent neural network (RNN) models: a classifier trained\non the joint vocabularies of all training languages; a Siamese RNN trained to\ndiscriminate between same and different words from multiple languages; and a\ncorrespondence autoencoder (CAE) RNN trained to reconstruct word pairs. In a\nword discrimination task on six target languages, all of these models\noutperform state-of-the-art unsupervised models trained on the zero-resource\nlanguages themselves, giving relative improvements of more than 30% in average\nprecision. When using only a few training languages, the multilingual CAE\nperforms better, but with more training languages the other multilingual models\nperform similarly. Using more training languages is generally beneficial, but\nimprovements are marginal on some languages. We present probing experiments\nwhich show that the CAE encodes more phonetic, word duration, language identity\nand speaker information than the other multilingual models.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC