This article describes a density ratio approach to integrating external\nLanguage Models (LMs) into end-to-end models for Automatic Speech Recognition\n(ASR). Applied to a Recurrent Neural Network Transducer (RNN-T) ASR model\ntrained on a given domain, a matched in-domain RNN-LM, and a target domain\nRNN-LM, the proposed method uses Bayes' Rule to define RNN-T posteriors for the\ntarget domain, in a manner directly analogous to the classic hybrid model for\nASR based on Deep Neural Networks (DNNs) or LSTMs in the Hidden Markov Model\n(HMM) framework (Bourlard & Morgan, 1994). The proposed approach is evaluated\nin cross-domain and limited-data scenarios, for which a significant amount of\ntarget domain text data is used for LM training, but only limited (or no)\n{audio, transcript} training data pairs are used to train the RNN-T.\nSpecifically, an RNN-T model trained on paired audio & transcript data from\nYouTube is evaluated for its ability to generalize to Voice Search data. The\nDensity Ratio method was found to consistently outperform the dominant approach\nto LM and end-to-end ASR integration, Shallow Fusion.\n
Paper
References (34)
Scroll for more · 22 remaining