Speech Emotion Recognition Based on BLSTM and CNN Feature Fusion

Speech emotion recognition (SER) is always challenging because of factors such as emotional corpus, acoustic features and SER modeling. SER based on deep learning are limited to using a spectrogram or handcrafted features as input, but cannot capture enough of the defects of emotional information, this paper proposes a feature fusion method based on Bidirectional Long Short-Term Memory (BLSTM) and Convolutional Neural Networks (CNN) to study richer emotional features, the method is combining context features and spatial features. Statistical features are used as the input of BLSTM network, the context features of speech signals are extracted by BLSTM, and the spatial features of speech signals are extracted by using log-mel spectrogram as the input of CNN, so as to jointly learn the emotional features with good recognition performance. The experimental results showed that the weighted accuracy and unweighted accuracy of the proposed method on the IEMOCAP data set were 74.14% and 65.62% respectively. In addition, compared with the existing SER methods, the effectiveness of the proposed method is verified.

Paper

Full text

PDF

Speech Emotion Recognition Based on BLSTM and CNN Feature Fusion

Semantic Scholar · Computer Science · 2020

Abstract

Speech emotion recognition (SER) is always challenging because of factors such as emotional corpus, acoustic features and SER modeling. SER based on deep learning are limited to using a spectrogram or handcrafted features as input, but cannot capture enough of the defects of emotional information, this paper proposes a feature fusion method based on Bidirectional Long Short-Term Memory (BLSTM) and Convolutional Neural Networks (CNN) to study richer emotional features, the method is combining context features and spatial features. Statistical features are used as the input of BLSTM network, the context features of speech signals are extracted by BLSTM, and the spatial features of speech signals are extracted by using log-mel spectrogram as the input of CNN, so as to jointly learn the emotional features with good recognition performance. The experimental results showed that the weighted accuracy and unweighted accuracy of the proposed method on the IEMOCAP data set were 74.14% and 65.62% respectively. In addition, compared with the existing SER methods, the effectiveness of the proposed method is verified.

Similar papers

© 2026 NYSGPT2525 LLC