Speech emotion recognition (SER) is always challenging because of factors such as emotional corpus, acoustic features and SER modeling. SER based on deep learning are limited to using a spectrogram or handcrafted features as input, but cannot capture enough of the defects of emotional information, this paper proposes a feature fusion method based on Bidirectional Long Short-Term Memory (BLSTM) and Convolutional Neural Networks (CNN) to study richer emotional features, the method is combining context features and spatial features. Statistical features are used as the input of BLSTM network, the context features of speech signals are extracted by BLSTM, and the spatial features of speech signals are extracted by using log-mel spectrogram as the input of CNN, so as to jointly learn the emotional features with good recognition performance. The experimental results showed that the weighted accuracy and unweighted accuracy of the proposed method on the IEMOCAP data set were 74.14% and 65.62% respectively. In addition, compared with the existing SER methods, the effectiveness of the proposed method is verified.
Paper
Full text
Speech Emotion Recognition Based on BLSTM and CNN Feature Fusion
Semantic Scholar · Computer Science · 2020
Abstract
Speech emotion recognition (SER) is always challenging because of factors such as emotional corpus, acoustic features and SER modeling. SER based on deep learning are limited to using a spectrogram or handcrafted features as input, but cannot capture enough of the defects of emotional information, this paper proposes a feature fusion method based on Bidirectional Long Short-Term Memory (BLSTM) and Convolutional Neural Networks (CNN) to study richer emotional features, the method is combining context features and spatial features. Statistical features are used as the input of BLSTM network, the context features of speech signals are extracted by BLSTM, and the spatial features of speech signals are extracted by using log-mel spectrogram as the input of CNN, so as to jointly learn the emotional features with good recognition performance. The experimental results showed that the weighted accuracy and unweighted accuracy of the proposed method on the IEMOCAP data set were 74.14% and 65.62% respectively. In addition, compared with the existing SER methods, the effectiveness of the proposed method is verified.