A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition

In this paper, we propose an efficient noise-robust approach to noisy speech emotion recognition (NSER). Although conventional NSER methods effectively handle some types of noise, such as stationary noise, they struggle with more complex noise encountered in realistic acoustic environments owing to its complexity and unpredictability. To address this issue, we introduce a novel NSER method that leverages automatic speech recognition (ASR) models as noise-robust feature extractors, filtering out non-vocal information from noisy speech. Specifically, we extract intermediate layer representations from the ASR model to capture emotional speech features for the NSER task. In this study, we conduct a comprehensive evaluation of ASR representations against traditional high-level statistical function (HSF) features, noise reduction approaches, mainstream NSER models, and fine-tuned self-supervised learning (SSL) methods. We analyze encoder and decoder groups through both single-layer and fusion-based multi-layer representations, further investigating the layer-wise performance and robustness of encoder and decoder modules under different noise types and intensities. Additionally, we explore their correlation with ASR performance, assess cross-modal performance using ASR transcripts, and examine cross-lingual generalization. Our experimental results demonstrate that (1) the proposed method outperforms HSF features, noise reduction approaches, mainstream NSER models, and SSL methods; (2) the adapter-based method further enhances ASR representations; (3) higher encoder and lower decoder layers (e.g., in Whisper) yield the best individual performance, while combining encoder and decoder layers across all depths achieves the strongest results; (4) performance decreases as noise intensity increases, especially with human speech noise; (5) this decline mirrors the degradation observed in ASR performance as noise intensity increases; (6) our approach surpasses transcript-based methods using ASR or ground-truth transcriptions; and (7) it achieves robust cross-lingual performance compared with mainstream SSL representations.

Paper

Similar papers

© 2026 NYSGPT2525 LLC