The existing face anti-spoofing (FAS) models have demonstrated high performance on specific datasets. However, for practical applications in real-world systems, it is essential to broaden the FAS model's capability to handle data from unknown domains, going beyond achieving strong results on a single baseline. Leveraging the remarkable capabilities of visual deformation models in discerning discriminative information, our exploration focuses on employing these models for recognizing facial presentation attacks in unexplored domains. To fulfill this objective, we introduce a groundbreaking Vision Transformer named DAA ViT, seamlessly integrating feature extraction and classification functionalities. Notably, we employ the Affine Consistent Module to fortify the geometric transformation stability. Simultaneously, the integration of a Deformable Attention module directs the model's focus towards crucial facial regions, thereby enhancing the extraction of FAS-related features. Our DAA ViT model demonstrates superior performance compared to contemporary techniques across publicly accessible face anti-spoofing datasets. We outperform existing methods in various aspects, including attack detection across diverse types, generalization proficiency, and other pertinent metrics. This substantiates the efficacy of our devised model framework and the integrated modules.
Paper
Full text
DAA-ViT: deformable attention affine-consistent vision transformer for face anti-spoofing
Semantic Scholar · Computer Science · 2024
Abstract
The existing face anti-spoofing (FAS) models have demonstrated high performance on specific datasets. However, for practical applications in real-world systems, it is essential to broaden the FAS model's capability to handle data from unknown domains, going beyond achieving strong results on a single baseline. Leveraging the remarkable capabilities of visual deformation models in discerning discriminative information, our exploration focuses on employing these models for recognizing facial presentation attacks in unexplored domains. To fulfill this objective, we introduce a groundbreaking Vision Transformer named DAA ViT, seamlessly integrating feature extraction and classification functionalities. Notably, we employ the Affine Consistent Module to fortify the geometric transformation stability. Simultaneously, the integration of a Deformable Attention module directs the model's focus towards crucial facial regions, thereby enhancing the extraction of FAS-related features. Our DAA ViT model demonstrates superior performance compared to contemporary techniques across publicly accessible face anti-spoofing datasets. We outperform existing methods in various aspects, including attack detection across diverse types, generalization proficiency, and other pertinent metrics. This substantiates the efficacy of our devised model framework and the integrated modules.