Deep Xi as a Front-End for Robust Automatic Speech Recognition

Currently, deep learning approaches to speech enhancement are most commonly used as front-ends for robust automatic speech recognition (ASR). A recently proposed deep learning approach to a priori SNR estimation, called Deep Xi, was able to produce enhanced speech at a higher quality and intelligibility than recent deep learning approaches to speech enhancement. Motivated by this, we investigate Deep Xi as a front-end for robust ASR. Deep Xi is evaluated using real-world non-stationary and coloured noise sources at multiple SNR levels. Our experimental investigation shows that Deep Xi as a frontend is able to produce a lower word error rate than recent deep learning approaches to speech enhancement. The results presented in this work show that Deep Xi is a viable front-end, and is able to significantly increase the robustness of an ASR system.Availability: Deep Xi is available at https://github.com/anicolson/DeepXi.

Paper

Similar papers

© 2026 NYSGPT2525 LLC