Self-Supervised Speaker Verification with Simple Siamese Network and Self-Supervised Regularization

Training speaker-discriminative and robust speaker verification systems\nwithout speaker labels is still challenging and worthwhile to explore. In this\nstudy, we propose an effective self-supervised learning framework and a novel\nregularization strategy to facilitate self-supervised speaker representation\nlearning. Different from contrastive learning-based self-supervised learning\nmethods, the proposed self-supervised regularization (SSReg) focuses\nexclusively on the similarity between the latent representations of positive\ndata pairs. We also explore the effectiveness of alternative online data\naugmentation strategies on both the time domain and frequency domain. With our\nstrong online data augmentation strategy, the proposed SSReg shows the\npotential of self-supervised learning without using negative pairs and it can\nsignificantly improve the performance of self-supervised speaker representation\nlearning with a simple Siamese network architecture. Comprehensive experiments\non the VoxCeleb datasets demonstrate that our proposed self-supervised approach\nobtains a 23.4% relative improvement by adding the effective self-supervised\nregularization and outperforms other previous works.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC