A comparison of self-supervised speech representations as input features for unsupervised acoustic word embeddings
Many speech processing tasks involve measuring the acoustic similarity\nbetween speech segments. Acoustic word embeddings (AWE) allow for efficient\ncomparisons by mapping speech segments of arbitrary duration to\nfixed-dimensional vectors. For zero-resource speech processing, where\nunlabelled speech is the only available resource, some of the best AWE\napproaches rely on weak top-down constraints in the form of automatically\ndiscovered word-like segments. Rather than learning embeddings at the segment\nlevel, another line of zero-resource research has looked at representation\nlearning at the short-time frame level. Recent approaches include\nself-supervised predictive coding and correspondence autoencoder (CAE) models.\nIn this paper we consider whether these frame-level features are beneficial\nwhen used as inputs for training to an unsupervised AWE model. We compare\nframe-level features from contrastive predictive coding (CPC), autoregressive\npredictive coding and a CAE to conventional MFCCs. These are used as inputs to\na recurrent CAE-based AWE model. In a word discrimination task on English and\nXitsonga data, all three representation learning approaches outperform MFCCs,\nwith CPC consistently showing the biggest improvement. In cross-lingual\nexperiments we find that CPC features trained on English can also be\ntransferred to Xitsonga.\n
Paper
References (51)
Scroll for more · 38 remaining