Semi-Supervised Speech Recognition via Graph-based Temporal Classification

Semi-supervised learning has demonstrated promising results in automatic\nspeech recognition (ASR) by self-training using a seed ASR model with\npseudo-labels generated for unlabeled data. The effectiveness of this approach\nlargely relies on the pseudo-label accuracy, for which typically only the\n1-best ASR hypothesis is used. However, alternative ASR hypotheses of an N-best\nlist can provide more accurate labels for an unlabeled speech utterance and\nalso reflect uncertainties of the seed ASR model. In this paper, we propose a\ngeneralized form of the connectionist temporal classification (CTC) objective\nthat accepts a graph representation of the training labels. The newly proposed\ngraph-based temporal classification (GTC) objective is applied for\nself-training with WFST-based supervision, which is generated from an N-best\nlist of pseudo-labels. In this setup, GTC is used to learn not only a temporal\nalignment, similarly to CTC, but also a label alignment to obtain the optimal\npseudo-label sequence from the weighted graph. Results show that this approach\ncan effectively exploit an N-best list of pseudo-labels with associated scores,\nconsiderably outperforming standard pseudo-labeling, with ASR results\napproaching an oracle experiment in which the best hypotheses of the N-best\nlists are selected manually.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC