Modern face recognition systems remain vulnerable to spoofing attempts, including both physical presentation attacks and digital forgeries. Traditionally, these two attacks vectors have been addressed by separate models or pipelines, each targeted to its specific artifacts and modalities. However, maintaining distinct detectors leads to increased system complexity, higher inference latency, and a combined attack vectors. We propose the Paired-Sampling Contrastive Framework, a unified training approach that leverages automatically matched pairs of genuine and attack selfies to learn modality-agnostic liveness clues. Evaluated on the 6th Face Anti-Spoofing Challenge “Unified Physical-Digital Attack Detection” benchmark, our method obtained an average classification error rate (ACER) of 2.10%, outperforming prior solutions. The proposed framework is lightweight, requires only 4.46 GFLOPs and a training runtime under one hour, making it practical for real-world deployment. Code and pretrained models are available at https://github.com/xPONYx/iccv2025_deepfake_challenge.
Paper
References (22)
Scroll for more · 10 remaining