Subband modeling for spoofing detection in automatic speaker\n verification

Spectrograms - time-frequency representations of audio signals - have found\nwidespread use in neural network-based spoofing detection. While deep models\nare trained on the fullband spectrum of the signal, we argue that not all\nfrequency bands are useful for these tasks. In this paper, we systematically\ninvestigate the impact of different subbands and their importance on replay\nspoofing detection on two benchmark datasets: ASVspoof 2017 v2.0 and ASVspoof\n2019 PA. We propose a joint subband modelling framework that employs n\ndifferent sub-networks to learn subband specific features. These are later\ncombined and passed to a classifier and the whole network weights are updated\nduring training. Our findings on the ASVspoof 2017 dataset suggest that the\nmost discriminative information appears to be in the first and the last 1 kHz\nfrequency bands, and the joint model trained on these two subbands shows the\nbest performance outperforming the baselines by a large margin. However, these\nfindings do not generalise on the ASVspoof 2019 PA dataset. This suggests that\nthe datasets available for training these models do not reflect real world\nreplay conditions suggesting a need for careful design of datasets for training\nreplay spoofing countermeasures.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC