A Comparison of Deep Learning and Machine Learning Models for Speech Emotion Recognition Using Multiple Features
This study presents a comparison between different deep learning and machine learning method for emotion recognition of speech data. Multiple features were extracted from the data and used for training the LSTM, MLP, Random Forest, and CNN models. Datasets used in the study are CREMA-D, RAVDESS, SAVEE, and TESS. The combined dataset are categorized into seven emotions such as surprise, neutral, disgust, fear, sad, happy, and angry. The features extracted from the dataset are zero crossing rate, MFCCs, root mean square, chroma stft, spectral contrast, spectral bandwidth, spectral centroid, spectral rolloff, spectral flatness, poly feature, mel spectrogram. LSTM model gives an accuracy of 57.5%. MLP model provided an accuracy of 59.6%. The accuracy of Random Forest model is 67.8% whereas the CNN model provided an accuracy of 75.8%. The results shows that CNN model out performs the LSTM, MLP, and Random Forest models when used for emotion recognition of speech.
Paper
Full text
A Comparison of Deep Learning and Machine Learning Models for Speech Emotion Recognition Using Multiple Features
Semantic Scholar · Computer Science · 2023
Abstract
This study presents a comparison between different deep learning and machine learning method for emotion recognition of speech data. Multiple features were extracted from the data and used for training the LSTM, MLP, Random Forest, and CNN models. Datasets used in the study are CREMA-D, RAVDESS, SAVEE, and TESS. The combined dataset are categorized into seven emotions such as surprise, neutral, disgust, fear, sad, happy, and angry. The features extracted from the dataset are zero crossing rate, MFCCs, root mean square, chroma stft, spectral contrast, spectral bandwidth, spectral centroid, spectral rolloff, spectral flatness, poly feature, mel spectrogram. LSTM model gives an accuracy of 57.5%. MLP model provided an accuracy of 59.6%. The accuracy of Random Forest model is 67.8% whereas the CNN model provided an accuracy of 75.8%. The results shows that CNN model out performs the LSTM, MLP, and Random Forest models when used for emotion recognition of speech.