Meta-learning with Latent Space Clustering in Generative Adversarial\n Network for Speaker Diarization

The performance of most speaker diarization systems with x-vector embeddings\nis both vulnerable to noisy environments and lacks domain robustness. Earlier\nwork on speaker diarization using generative adversarial network (GAN) with an\nencoder network (ClusterGAN) to project input x-vectors into a latent space has\nshown promising performance on meeting data. In this paper, we extend the\nClusterGAN network to improve diarization robustness and enable rapid\ngeneralization across various challenging domains. To this end, we fetch the\npre-trained encoder from the ClusterGAN and fine-tune it by using prototypical\nloss (meta-ClusterGAN or MCGAN) under the meta-learning paradigm. Experiments\nare conducted on CALLHOME telephonic conversations, AMI meeting data, DIHARD II\n(dev set) which includes challenging multi-domain corpus, and two\nchild-clinician interaction corpora (ADOS, BOSCC) related to the autism\nspectrum disorder domain. Extensive analyses of the experimental data are done\nto investigate the effectiveness of the proposed ClusterGAN and MCGAN\nembeddings over x-vectors. The results show that the proposed embeddings with\nnormalized maximum eigengap spectral clustering (NME-SC) back-end consistently\noutperform Kaldi state-of-the-art z-vector diarization system. Finally, we\nemploy embedding fusion with x-vectors to provide further improvement in\ndiarization performance. We achieve a relative diarization error rate (DER)\nimprovement of 6.67% to 53.93% on the aforementioned datasets using the\nproposed fused embeddings over x-vectors. Besides, the MCGAN embeddings provide\nbetter performance in the number of speakers estimation and short speech\nsegment diarization as compared to x-vectors and ClusterGAN in telephonic data.\n

Paper

References (76)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC