Automated Speech Emotion Recognition on Smart Phones

The emergence of Speech Emotion Recognition (SER) as a focal point of research into speech processing reflects its significance in the field of Human-Computer Interaction (HCI). It is core required functionality for a variety of applications and a high degree of accuracy is critical; activities as diverse as evaluating levels of emotion in children in care and measuring customer satisfaction. The extent of demand for accurate SER is reflected in the significant number of papers that have been published and studies performed. An innovative approach to speech recognition is presented in this paper centered on a cloud model alongside the conventional system for measuring emotion in speech. There are multiple stages involved in detecting and identifying emotions in speech from audio clips. The initial pre-processing stage detects the speech in the audio file and applies noise reduction. Next, the system uses Mel-frequency cepstral coefficient (MFCC) algorithms to extract features.This process results in the creation of testing and training datasets populated with the emotions: Neutral, Happiness, Sadness, Fear, Surprise, Disgust and Anger. The classification stage utilizes Support Vector Machine (SVM) classifiers to identify the emotion. An additional step implements a Confusion Matrix (CM) method to assess how these classifiers performed. Testing was executed against RAVDESS and SAVEE databases, where the detection rate achieved against the RAVDESS database was 95.3%.

Paper

Full text

PDF

Automated Speech Emotion Recognition on Smart Phones

Semantic Scholar · Computer Science · 2018

Abstract

The emergence of Speech Emotion Recognition (SER) as a focal point of research into speech processing reflects its significance in the field of Human-Computer Interaction (HCI). It is core required functionality for a variety of applications and a high degree of accuracy is critical; activities as diverse as evaluating levels of emotion in children in care and measuring customer satisfaction. The extent of demand for accurate SER is reflected in the significant number of papers that have been published and studies performed. An innovative approach to speech recognition is presented in this paper centered on a cloud model alongside the conventional system for measuring emotion in speech. There are multiple stages involved in detecting and identifying emotions in speech from audio clips. The initial pre-processing stage detects the speech in the audio file and applies noise reduction. Next, the system uses Mel-frequency cepstral coefficient (MFCC) algorithms to extract features.This process results in the creation of testing and training datasets populated with the emotions: Neutral, Happiness, Sadness, Fear, Surprise, Disgust and Anger. The classification stage utilizes Support Vector Machine (SVM) classifiers to identify the emotion. An additional step implements a Confusion Matrix (CM) method to assess how these classifiers performed. Testing was executed against RAVDESS and SAVEE databases, where the detection rate achieved against the RAVDESS database was 95.3%.

Similar papers

© 2026 NYSGPT2525 LLC