Robust Text-Dependent Speaker Verification via Character-Level\n Information Preservation for the SdSV Challenge 2020

This paper describes our submission to Task 1 of the Short-duration Speaker\nVerification (SdSV) challenge 2020. Task 1 is a text-dependent speaker\nverification task, where both the speaker and phrase are required to be\nverified. The submitted systems were composed of TDNN-based and ResNet-based\nfront-end architectures, in which the frame-level features were aggregated with\nvarious pooling methods (e.g., statistical, self-attentive, ghostVLAD pooling).\nAlthough the conventional pooling methods provide embeddings with a sufficient\namount of speaker-dependent information, our experiments show that these\nembeddings often lack phrase-dependent information. To mitigate this problem,\nwe propose a new pooling and score compensation methods that leverage a\nCTC-based automatic speech recognition (ASR) model for taking the lexical\ncontent into account. Both methods showed improvement over the conventional\ntechniques, and the best performance was achieved by fusing all the\nexperimented systems, which showed 0.0785% MinDCF and 2.23% EER on the\nchallenge's evaluation subset.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC