On Scaling Contrastive Representations for Low-Resource Speech Recognition

Recent advances in self-supervised learning through contrastive training have\nshown that it is possible to learn a competitive speech recognition system with\nas little as 10 minutes of labeled data. However, these systems are\ncomputationally expensive since they require pre-training followed by\nfine-tuning in a large parameter space. We explore the performance of such\nsystems without fine-tuning by training a state-of-the-art speech recognizer on\nthe fixed representations from the computationally demanding wav2vec 2.0\nframework. We find performance to decrease without fine-tuning and, in the\nextreme low-resource setting, wav2vec 2.0 is inferior to its predecessor. In\naddition, we find that wav2vec 2.0 representations live in a low dimensional\nsubspace and that decorrelating the features of the representations can\nstabilize training of the automatic speech recognizer. Finally, we propose a\nbidirectional extension to the original wav2vec framework that consistently\nimproves performance.\n

Paper

References (21)

Scroll for more · 9 remaining

Similar papers

© 2026 NYSGPT2525 LLC