Annotators' agreement and spontaneous emotion classification performance

Abstract The combination of various types of data can significantly in-crease the amount of emotional material for training of morereliable real-life emotion classifiers. There are two well-knownschemes of annotation utilized for emotional speech: multi-dimensional and categories-based. Multi-dimensional anno-tation is usually applied for labeling spontaneous emotionalevents, and categorial-based annotation is used for specifica-tion of the acted ”full blown” emotional chunks. In order tosimulate real-life conditions we used a cross-corpora evalua-tion strategy for datasets with different schemes of emotionalannotation. Emotional models were trained on acted materialfrom the EMO-DB (categories based annotation) dataset andevaluated on spontaneous data from the VAM dataset (multi-dimensional annotation). The best emotion classification per-formance was obtained on real-life emotional instances withthe most intense arousal labels provided by a majority votingstrategy (out of 17 annotators). We find that the correspond-ing spontaneous speech samples containing the most intensiveemotional content are comparable with acted instances. The im-portance of employing a larger number of emotional annotatorswas finally addressed in our article.Index Terms: emotion recognition, cross-corpora evaluation,phoneme-level emotional models, turn-level emotional models,emotional intensity

Paper

Full text

PDF

Annotators' agreement and spontaneous emotion classification performance

Semantic Scholar · Computer Science · 2015

Abstract

Abstract The combination of various types of data can significantly in-crease the amount of emotional material for training of morereliable real-life emotion classifiers. There are two well-knownschemes of annotation utilized for emotional speech: multi-dimensional and categories-based. Multi-dimensional anno-tation is usually applied for labeling spontaneous emotionalevents, and categorial-based annotation is used for specifica-tion of the acted ”full blown” emotional chunks. In order tosimulate real-life conditions we used a cross-corpora evalua-tion strategy for datasets with different schemes of emotionalannotation. Emotional models were trained on acted materialfrom the EMO-DB (categories based annotation) dataset andevaluated on spontaneous data from the VAM dataset (multi-dimensional annotation). The best emotion classification per-formance was obtained on real-life emotional instances withthe most intense arousal labels provided by a majority votingstrategy (out of 17 annotators). We find that the correspond-ing spontaneous speech samples containing the most intensiveemotional content are comparable with acted instances. The im-portance of employing a larger number of emotional annotatorswas finally addressed in our article.Index Terms: emotion recognition, cross-corpora evaluation,phoneme-level emotional models, turn-level emotional models,emotional intensity

Similar papers

© 2026 NYSGPT2525 LLC