Disinformation increasingly spans multiple modalities, including manipulated audio, fake videos, and text-image content pairs in the case of articles. Existing detection models often address the modalities separately, limiting their effectiveness in real-world scenarios. This study proposes a unified multimodal disinformation detection model that simultaneously analyzes text, image, and video content. To operationalize this unified approach, we transform video data into complementary textual and visual representations. Audio tracks are transcribed using Whisper, while keyframes are extracted from video using one of three methods: random frame extraction, clustering-based selection, and our novel extraction method. Captions are generated for each keyframe to embed visual semantics into the textual stream, enabling integrated cross-modal analysis. This combined representation is evaluated against unimodal baselines and state-of-the-art Vision-Language Models (VLMs), including LLaMA and VILA. Results across model architectures and dataset configurations show that our unified multimodal pipeline outperforms separate modality-specific systems in detecting disinformation.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex