In this work, we investigate if the wav2vec 2.0 self-supervised pretraining\nhelps mitigate the overfitting issues with connectionist temporal\nclassification (CTC) training to reduce its performance gap with flat-start\nlattice-free MMI (E2E-LFMMI) for automatic speech recognition with limited\ntraining data. Towards that objective, we use the pretrained wav2vec 2.0 BASE\nmodel and fine-tune it on three different datasets including out-of-domain\n(Switchboard) and cross-lingual (Babel) scenarios. Our results show that for\nsupervised adaptation of the wav2vec 2.0 model, both E2E-LFMMI and CTC achieve\nsimilar results; significantly outperforming the baselines trained only with\nsupervised data. Fine-tuning the wav2vec 2.0 model with E2E-LFMMI and CTC we\nobtain the following relative WER improvements over the supervised baseline\ntrained with E2E-LFMMI. We get relative improvements of 40% and 44% on the\nclean-set and 64% and 58% on the test set of Librispeech (100h) respectively.\nOn Switchboard (300h) we obtain relative improvements of 33% and 35%\nrespectively. Finally, for Babel languages, we obtain relative improvements of\n26% and 23% on Swahili (38h) and 18% and 17% on Tagalog (84h) respectively.\n
Paper
References (26)
Scroll for more · 14 remaining