Joint Optimization of Speaker and Spoof Detectors for Spoofing-Robust Automatic Speaker Verification
Spoofing-robust speaker verification (SASV) combines the tasks of speaker and spoof detection to authenticate speakers under adversarial settings. Many SASV systems rely on fusion of speaker and spoof cues based on independently trained subsystems, which often limits joint performance. In this study, we propose a novel modular, yet jointly optimized, SASV framework that integrates the outputs of speaker and spoofing detection subsystems using trainable back-end classifiers. Our framework enables direct optimization of both subsystems under a unified objective, using the recently-proposed architecture-agnostic detection cost function (a-DCF) as the training objective. This approach preserves the interpretability and plug-and-play compatibility of standalone detectors while aligning them towards a common goal. Our experiments on the ASVspoof 5 dataset demonstrate two important findings: (i) nonlinear score fusion consistently improves a-DCF over linear fusion, and (ii) the combination of weighted cosine scoring for speaker detection with SSL-AASIST for spoof detection achieves state-of-the-art performance, reducing min a-DCF to 0.196 and SPF-EER to 7.6%. These contributions highlight the importance of modular design, calibrated integration, and task-aligned optimization for advancing robust and interpretable SASV systems.