FEENN: The Feature Enhancement Embedded Neural Network for Robust Multimodal Emotion Recognition
This paper introduces the Delta Team's submission to the Multimodal Emotion Recognition(MER 2023)-MER-NOISE Challenge. Multimodal emotion recognition aims to improve sentiment analysis ability by integrating emotional information from multiple modalities. However, in real applications, various interference such as noise and complex backgrounds significantly degrade model performance, making model construction more challenging. In this paper, we propose two simple and effective feature enhancement strategies, Hidden State Smoothing (HSS) and Embedding Vector Augmentation (EVA), to improve the robustness of the model. The HSS strategy extracts text and audio feature vectors by weighting encoded features at both shallow and deep hidden states, enhancing the model's representation of fine-grained sentiment categories. Additionally, the EVA is introduced to perturb audio embedding vectors with random noise for improving model generalization ability. On the MER-NOISE track, the models are trained with limited video data, and both strategies effectively improve the robustness of the model to varying degrees. The HSS strategy proves to be particularly effective and concise, achieving a 30% performance improvement over the baseline on the noisy test set. It adapts to many models in various interference scenarios.
Paper
Full text
FEENN: The Feature Enhancement Embedded Neural Network for Robust Multimodal Emotion Recognition
Semantic Scholar · Computer Science · 2023
Abstract
This paper introduces the Delta Team's submission to the Multimodal Emotion Recognition(MER 2023)-MER-NOISE Challenge. Multimodal emotion recognition aims to improve sentiment analysis ability by integrating emotional information from multiple modalities. However, in real applications, various interference such as noise and complex backgrounds significantly degrade model performance, making model construction more challenging. In this paper, we propose two simple and effective feature enhancement strategies, Hidden State Smoothing (HSS) and Embedding Vector Augmentation (EVA), to improve the robustness of the model. The HSS strategy extracts text and audio feature vectors by weighting encoded features at both shallow and deep hidden states, enhancing the model's representation of fine-grained sentiment categories. Additionally, the EVA is introduced to perturb audio embedding vectors with random noise for improving model generalization ability. On the MER-NOISE track, the models are trained with limited video data, and both strategies effectively improve the robustness of the model to varying degrees. The HSS strategy proves to be particularly effective and concise, achieving a 30% performance improvement over the baseline on the noisy test set. It adapts to many models in various interference scenarios.