Vector-quantized neural networks for acoustic unit discovery in the ZeroSpeech 2020 challenge

In this paper, we explore vector quantization for acoustic unit discovery.\nLeveraging unlabelled data, we aim to learn discrete representations of speech\nthat separate phonetic content from speaker-specific details. We propose two\nneural models to tackle this challenge - both use vector quantization to map\ncontinuous features to a finite set of codes. The first model is a type of\nvector-quantized variational autoencoder (VQ-VAE). The VQ-VAE encodes speech\ninto a sequence of discrete units before reconstructing the audio waveform. Our\nsecond model combines vector quantization with contrastive predictive coding\n(VQ-CPC). The idea is to learn a representation of speech by predicting future\nacoustic units. We evaluate the models on English and Indonesian data for the\nZeroSpeech 2020 challenge. In ABX phone discrimination tests, both models\noutperform all submissions to the 2019 and 2020 challenges, with a relative\nimprovement of more than 30%. The models also perform competitively on a\ndownstream voice conversion task. Of the two, VQ-CPC performs slightly better\nin general and is simpler and faster to train. Finally, probing experiments\nshow that vector quantization is an effective bottleneck, forcing the models to\ndiscard speaker information.\n

Paper

References (43)

Scroll for more · 31 remaining

Similar papers

© 2026 NYSGPT2525 LLC