Emotion recognition from audio signal requires feature extraction and classifier training. The feature vector consists of elements of the audio signal which characterise speaker specific features such as tone, pitch, energy, which is crucial to train the classifier model to recognise a particular emotion accurately. The North American English language open source dataset was divided into training and testing manually. Speaker vocal tract information, represented by Mel-frequency cepstral coefficients (MFCC), was extracted from the audio samples in training dataset. Pitch, Short Term Energy(STE), and MFCC coefficients of audio samples in emotions anger, happiness, and sadness were obtained. These extracted feature vectors were sent to the classifier model. The test dataset will undergo the extraction procedure following which the classifier would make a decision regarding the underlying emotion in the test audio. The training and test databases used were North American English acted and natural speech corpus, real-time input English speech, regional language databases in Hindi and Marathi. The paper details the two methods applied on feature vectors and the effect of increasing the number of feature vectors fed to the classifier. It provides an analysis of the accuracy of classification for Indian English speech and speech in Hindi and Marathi. The achieved accuracy for Indian English speech was 80 percent.
Paper
Full text
Speech based Emotion Recognition using Machine Learning
Semantic Scholar · Computer Science · 2019
Abstract
Emotion recognition from audio signal requires feature extraction and classifier training. The feature vector consists of elements of the audio signal which characterise speaker specific features such as tone, pitch, energy, which is crucial to train the classifier model to recognise a particular emotion accurately. The North American English language open source dataset was divided into training and testing manually. Speaker vocal tract information, represented by Mel-frequency cepstral coefficients (MFCC), was extracted from the audio samples in training dataset. Pitch, Short Term Energy(STE), and MFCC coefficients of audio samples in emotions anger, happiness, and sadness were obtained. These extracted feature vectors were sent to the classifier model. The test dataset will undergo the extraction procedure following which the classifier would make a decision regarding the underlying emotion in the test audio. The training and test databases used were North American English acted and natural speech corpus, real-time input English speech, regional language databases in Hindi and Marathi. The paper details the two methods applied on feature vectors and the effect of increasing the number of feature vectors fed to the classifier. It provides an analysis of the accuracy of classification for Indian English speech and speech in Hindi and Marathi. The achieved accuracy for Indian English speech was 80 percent.