Unsupervised Representation Learning for Speaker Recognition via\n Contrastive Equilibrium Learning

In this paper, we propose a simple but powerful unsupervised learning method\nfor speaker recognition, namely Contrastive Equilibrium Learning (CEL), which\nincreases the uncertainty on nuisance factors latent in the embeddings by\nemploying the uniformity loss. Also, to preserve speaker discriminability, a\ncontrastive similarity loss function is used together. Experimental results\nshowed that the proposed CEL significantly outperforms the state-of-the-art\nunsupervised speaker verification systems and the best performing model\nachieved 8.01% and 4.01% EER on VoxCeleb1 and VOiCES evaluation sets,\nrespectively. On top of that, the performance of the supervised speaker\nembedding networks trained with initial parameters pre-trained via CEL showed\nbetter performance than those trained with randomly initialized parameters.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC