FEENN: The Feature Enhancement Embedded Neural Network for Robust Multimodal Emotion Recognition

This paper introduces the Delta Team's submission to the Multimodal Emotion Recognition(MER 2023)-MER-NOISE Challenge. Multimodal emotion recognition aims to improve sentiment analysis ability by integrating emotional information from multiple modalities. However, in real applications, various interference such as noise and complex backgrounds significantly degrade model performance, making model construction more challenging. In this paper, we propose two simple and effective feature enhancement strategies, Hidden State Smoothing (HSS) and Embedding Vector Augmentation (EVA), to improve the robustness of the model. The HSS strategy extracts text and audio feature vectors by weighting encoded features at both shallow and deep hidden states, enhancing the model's representation of fine-grained sentiment categories. Additionally, the EVA is introduced to perturb audio embedding vectors with random noise for improving model generalization ability. On the MER-NOISE track, the models are trained with limited video data, and both strategies effectively improve the robustness of the model to varying degrees. The HSS strategy proves to be particularly effective and concise, achieving a 30% performance improvement over the baseline on the noisy test set. It adapts to many models in various interference scenarios.

Paper

Full text

PDF

FEENN: The Feature Enhancement Embedded Neural Network for Robust Multimodal Emotion Recognition

Semantic Scholar · Computer Science · 2023

Abstract

This paper introduces the Delta Team's submission to the Multimodal Emotion Recognition(MER 2023)-MER-NOISE Challenge. Multimodal emotion recognition aims to improve sentiment analysis ability by integrating emotional information from multiple modalities. However, in real applications, various interference such as noise and complex backgrounds significantly degrade model performance, making model construction more challenging. In this paper, we propose two simple and effective feature enhancement strategies, Hidden State Smoothing (HSS) and Embedding Vector Augmentation (EVA), to improve the robustness of the model. The HSS strategy extracts text and audio feature vectors by weighting encoded features at both shallow and deep hidden states, enhancing the model's representation of fine-grained sentiment categories. Additionally, the EVA is introduced to perturb audio embedding vectors with random noise for improving model generalization ability. On the MER-NOISE track, the models are trained with limited video data, and both strategies effectively improve the robustness of the model to varying degrees. The HSS strategy proves to be particularly effective and concise, achieving a 30% performance improvement over the baseline on the noisy test set. It adapts to many models in various interference scenarios.

Similar papers

© 2026 NYSGPT2525 LLC