Multimodal Emotion Recognition and Sentiment Analysis Using Masked Attention and Multimodal Interaction

People express emotions verbally (with the linguistic part) and non-verbally (with facial expressions and speech tone). For better emotion recognition it is expedient to use both types of expressions. In this paper, a multimodal approach for emotion recognition based on fusion of textual, audio, and video data with masked multimodal attention and multimodal interaction are suggested. Our models is built on top of the language representation model Bidirectional Encoder Representations from Transformers (BERT). BERT is mostly used to work with text data, while the approaches with the interaction of text and audio modalities with fine-tuning a pre-trained BERT model are less common. In this work, a new 3-Modal Cross-BERT model that utilizes BERT fine-tuning based on textual, audio, and video data using masked multimodal attention and the model with multimodal interaction are proposed. Our algorithms were evaluated on publicly available multimodal sentiment and emotion analysis datasets CMU-MOSI, CMU-MOSEI, IEMOCAP and MELD. Experimental results show significant improvements in the performance across all metrics compared to the previous state-of-the-art methods for chosen datasets.

Paper

Full text

PDF

Multimodal Emotion Recognition and Sentiment Analysis Using Masked Attention and Multimodal Interaction

Semantic Scholar · Computer Science · 2023

Abstract

People express emotions verbally (with the linguistic part) and non-verbally (with facial expressions and speech tone). For better emotion recognition it is expedient to use both types of expressions. In this paper, a multimodal approach for emotion recognition based on fusion of textual, audio, and video data with masked multimodal attention and multimodal interaction are suggested. Our models is built on top of the language representation model Bidirectional Encoder Representations from Transformers (BERT). BERT is mostly used to work with text data, while the approaches with the interaction of text and audio modalities with fine-tuning a pre-trained BERT model are less common. In this work, a new 3-Modal Cross-BERT model that utilizes BERT fine-tuning based on textual, audio, and video data using masked multimodal attention and the model with multimodal interaction are proposed. Our algorithms were evaluated on publicly available multimodal sentiment and emotion analysis datasets CMU-MOSI, CMU-MOSEI, IEMOCAP and MELD. Experimental results show significant improvements in the performance across all metrics compared to the previous state-of-the-art methods for chosen datasets.

Similar papers

© 2026 NYSGPT2525 LLC