Improving Speaker Verification Robustness With Multilingual Phonetic Information and Feature Decorrelation

In automatic speaker verification (ASV) systems, the representational capability of speaker embeddings is critical to system performance. Existing studies have shown that introducing phonetic information during speaker embedding extraction can achieve phonetic alignment, thereby enhancing discriminative ability. However, current methods typically rely on monolingual phonetic information, which limits the effectiveness of phonetic alignment when training models on multilingual ASV datasets. Moreover, as the parameters of the ASV model increase, the extracted speaker embeddings may include both speaker-related features and speaker-independent features influenced by the training data. The correlations among these features can increase the risk of overfitting in the model. To address these challenges, this paper proposes a method that combines multilingual phonetic information (MuPI) and feature decorrelation (FD), termed MuPI-FD. Specifically, we analyze the language distribution in the training data and train three multilingual automatic speech recognition (ASR) models with different architectures. Using transfer learning, we incorporate multilingual phonetic information into ASV systems to achieve frame-level phonetic alignment. Additionally, to ensure that the extracted speaker embeddings focus more on speaker-related features, we introduce Random Fourier Features (RFF) and a sample reweighting method to reduce correlations among speaker embedding features. Experimental results demonstrate that ASV systems introducing multilingual phonetic information significantly outperform those using monolingual phonetic information on the VoxCeleb, SITW, and CN-Celeb datasets. Furthermore, ASV systems processed with feature decorrelation exhibit stronger generalization capabilities. Compared to the MFA-Conformer baseline system, MuPI-FD achieves an average relative improvement of 31.89% in EER across all test sets.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at Semantic Scholar

Similar papers

© 2026 NYSGPT2525 LLC