With the continuous development of speech recognition technology, speaker verification (SV) has become an important method for identity authentication. Traditional SV methods rely on handcrafted feature extraction, while the introduction of deep learning has significantly improved system performance. However, the scarcity of labeled data still limits the widespread application of deep learning methods in SV. Self-supervised learning, by mining the latent information in massive unlabeled data, enhances the model’s generalization ability and has become a key technology to address this issue.DINO is an efficient self-supervised learning method that generates pseudo-labels from unlabeled speech data through clustering algorithms, providing support for subsequent training. However, the clustering process may produce noisy pseudolabels, which can reduce the overall recognition performance of the system and restrict further improvement of the model’s performance.To address this issue, this paper proposes an improved clustering framework based on similarity connection graphs and Graph Convolutional Networks (GCN). By leveraging GCN’s strength in modeling structured data and incorporating the relational information between nodes in the similarity connection graph, the clustering process is optimized, improving the accuracy of pseudo-labels and thereby enhancing the robustness and performance of the self-supervised speaker verification system. Experimental results show that this method can significantly improve system performance and provide a new approach for self-supervised speaker verification.
Paper
References (31)
Scroll for more · 19 remaining