M3FEND: Multi-Modal Mixture of Experts with Adversarial Gating for Multi-Modal Fake News Detection
Typical approaches in Multi-modal Fake News Detection (MFND) often utilize cross-modal Transformers to capture semantic interactions between textual and visual modalities. However, traditional methods process visual semantics on a holistic scale, ignoring composite semantic information on elemental scales, resulting in limited model performance. To address this, we propose the Multi-modal Mixture-of-Experts with Adversarial Gating for Multi-modal Fake News Detection (M3FEND). Firstly, M3FEND transforms the visual modality at elemental scales to generate the visual semantic abstract and optical characters. Furthermore, M3FEND trains the Multi-modal Fusion Experts through fusion paths tailored to different scales, enabling comprehensive fusion between visual and textual semantics. Then, we propose the Dynamic Scoring Gating Network to reduce noise scale interference. By analyzing the gradient influence of each scale, we estimate each scale’s importance to generate gating weights that reduce noise introduced during the fusion process. Experiments on two datasets demonstrate that M3FEND outperforms state-of-the-art methods.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex