Investigating Methods to Improve Language Model Integration for Attention-based Encoder-Decoder ASR Models
Attention-based encoder-decoder (AED) models learn an implicit internal\nlanguage model (ILM) from the training transcriptions. The integration with an\nexternal LM trained on much more unpaired text usually leads to better\nperformance. A Bayesian interpretation as in the hybrid autoregressive\ntransducer (HAT) suggests dividing by the prior of the discriminative acoustic\nmodel, which corresponds to this implicit LM, similarly as in the hybrid hidden\nMarkov model approach. The implicit LM cannot be calculated efficiently in\ngeneral and it is yet unclear what are the best methods to estimate it. In this\nwork, we compare different approaches from the literature and propose several\nnovel methods to estimate the ILM directly from the AED model. Our proposed\nmethods outperform all previous approaches. We also investigate other methods\nto suppress the ILM mainly by decreasing the capacity of the AED model,\nlimiting the label context, and also by training the AED model together with a\npre-existing LM.\n
Paper
References (26)
Scroll for more · 14 remaining