Deep Architecture Enhancing Robustness to Noise, Adversarial Attacks,\n and Cross-corpus Setting for Speech Emotion Recognition

Speech emotion recognition systems (SER) can achieve high accuracy when the\ntraining and test data are identically distributed, but this assumption is\nfrequently violated in practice and the performance of SER systems plummet\nagainst unforeseen data shifts. The design of robust models for accurate SER is\nchallenging, which limits its use in practical applications. In this paper we\npropose a deeper neural network architecture wherein we fuse DenseNet, LSTM and\nHighway Network to learn powerful discriminative features which are robust to\nnoise. We also propose data augmentation with our network architecture to\nfurther improve the robustness. We comprehensively evaluate the architecture\ncoupled with data augmentation against (1) noise, (2) adversarial attacks and\n(3) cross-corpus settings. Our evaluations on the widely used IEMOCAP and\nMSP-IMPROV datasets show promising results when compared with existing studies\nand state-of-the-art models.\n

Paper

References (43)

Scroll for more · 31 remaining

Similar papers

© 2026 NYSGPT2525 LLC