A Comprehensive Deepfake Detection Framework Based on Multimodal Liveness Detection and Deep Learning Integration

Deepfake technology poses severe security and trust challenges, demanding robust detection strategies.We propose a unified audio-visual framework that jointly exploits facial dynamics and speech acoustics to capture subtle forgery traces.The system integrates CNN-based face encoders for spatial features, RNN/Temporal-Conformer blocks for temporal cues, and lightweight Transformers for contextual modeling.An adaptive mid-late fusion module aggregates multimodal embeddings via gated attention and a calibration head, ensuring resilience against partial modality corruption.Preprocessing involves face detection and cropping, log-mel spectrogram extraction from 16 kHz audio, and alignment of video-audio segments with a unimodal fallback mechanism for instances of missing modalities.Experiments on FaceForensics++, DFDC, and FakeAVCeleb-with subject-and speakerdisjoint splits-demonstrate strong generalization, achieving up to 97.3% accuracy and 99.0%AUC under clean conditions.Key contributions include a principled multimodal fusion strategy and a plug-and-play ensemble mechanism that stabilizes training across datasets.Limitations such as computational overhead and the need for further adversarial robustness testing are discussed, while reproducibility is facilitated through released configurations and scripts.Overall, this work advances multimodal Deepfake detection by offering an efficient, accurate, and extensible defense framework.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC