Knowledge Distillation and Data Selection for Semi-Supervised Learning in CTC Acoustic Models

Semi-supervised learning (SSL) is an active area of research which aims to\nutilize unlabelled data in order to improve the accuracy of speech recognition\nsystems. The current study proposes a methodology for integration of two key\nideas: 1) SSL using connectionist temporal classification (CTC) objective and\nteacher-student based learning 2) Designing effective data-selection mechanisms\nfor leveraging unlabelled data to boost performance of student models. Our aim\nis to establish the importance of good criteria in selecting samples from a\nlarge pool of unlabelled data based on attributes like confidence measure,\nspeaker and content variability. The question we try to answer is: Is it\npossible to design a data selection mechanism which reduces dependence on a\nlarge set of randomly selected unlabelled samples without compromising on Word\nError Rate (WER)? We perform empirical investigations of different data\nselection methods to answer this question and quantify the effect of different\nsampling strategies. On a semi-supervised ASR setting with 40000 hours of\ncarefully selected unlabelled data, our CTC-SSL approach gives 17% relative WER\nimprovement over a baseline CTC system trained with labelled data. It also\nachieves on-par performance with CTC-SSL system trained on order of magnitude\nlarger unlabeled data based on random sampling.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC