Multimodal sentiment analysis aims to extract and integrate semantic\ninformation collected from multiple modalities to recognize the expressed\nemotions and sentiment in multimodal data. This research area's major concern\nlies in developing an extraordinary fusion scheme that can extract and\nintegrate key information from various modalities. However, one issue that may\nrestrict previous work to achieve a higher level is the lack of proper modeling\nfor the dynamics of the competition between the independence and relevance\namong modalities, which could deteriorate fusion outcomes by causing the\ncollapse of modality-specific feature space or introducing extra noise. To\nmitigate this, we propose the Bi-Bimodal Fusion Network (BBFN), a novel\nend-to-end network that performs fusion (relevance increment) and separation\n(difference increment) on pairwise modality representations. The two parts are\ntrained simultaneously such that the combat between them is simulated. The\nmodel takes two bimodal pairs as input due to the known information imbalance\namong modalities. In addition, we leverage a gated control mechanism in the\nTransformer architecture to further improve the final output. Experimental\nresults on three datasets (CMU-MOSI, CMU-MOSEI, and UR-FUNNY) verifies that our\nmodel significantly outperforms the SOTA. The implementation of this work is\navailable at https://github.com/declare-lab/multimodal-deep-learning.\n
Paper
References (50)
Scroll for more · 38 remaining