EmoNet: A Transfer Learning Framework for Multi-Corpus Speech Emotion Recognition

In this manuscript, the topic of multi-corpus Speech Emotion Recognition (SER) is approached from a deep transfer learning perspective. A large corpus of emotional speech data, <bold><sc>EmoSet</sc></bold>, is assembled from a number of existing Speech Emotion Recognition (SER) corpora. In total, <sc>EmoSet</sc> contains <bold>84 181 audio recordings</bold> from <bold>26 SER corpora</bold> with a total duration of over <bold>65 hours</bold>. The corpus is then utilised to create a novel framework for multi-corpus SER and general audio recognition, namely <bold><sc>EmoNet</sc></bold>. A combination of a deep ResNet architecture and residual adapters is transferred from the field of multi-domain visual recognition to multi-corpus SER on <sc>EmoSet</sc>. The introduced residual adapter approach enables parameter efficient training of a multi-domain SER model on all 26 corpora. A shared model with only 3.5 times the number of parameters of a model trained on a single database leads to increased performance for 21 of the 26 corpora in <sc>EmoSet</sc>. Using repeated training runs and Almost Stochastic Order with significance level of <inline-formula><tex-math notation="LaTeX">$\alpha = 0.05$</tex-math><alternatives><mml:math><mml:mrow><mml:mi>α</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>05</mml:mn></mml:mrow></mml:math><inline-graphic xlink:href="gerczuk-ieq1-3135152.gif"/></alternatives></inline-formula>, these improvements are further significant for 15 datasets while there are just three corpora that see only significant decreases across the residual adapter transfer experiments. Finally, we make our <sc>EmoNet</sc> framework publicly available for users and developers at <monospace><uri>https://github.com/EIHW/EmoNet</uri></monospace>.

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC