Multi-level Data Filter for Speaker Recognition on Imbalanced Datasets

End to end speaker embedding systems have shown promising performance on speaker recognition tasks. However, speaker recognition system often degrade obviously when there are different scale and discrepant distribution of training datasets. In this paper, we proposed a novel model comprised of MultiReader technique and ResNet-GhostVLAD network, which makes the performance on speaker recognition task more stable and excellent. The MultiReader technique and dictionary-based GhostVLAD layer enable our model to filter data at data-level and feature-level respectively, and make it robust for imbalanced datasets which contain noisy and irrelevant information. The proposed model has been trained on the VoxCeleb and AISHELL combined dataset, then tested on AISHELL dataset. Evaluations for different weights of the multiple datasets show that our model outperforms the method of directly mixing training set by obvious margins, which are 10.7% and 44.3% relative improvement.

Paper

Full text

PDF

Multi-level Data Filter for Speaker Recognition on Imbalanced Datasets

Semantic Scholar · Computer Science · 2019

Abstract

End to end speaker embedding systems have shown promising performance on speaker recognition tasks. However, speaker recognition system often degrade obviously when there are different scale and discrepant distribution of training datasets. In this paper, we proposed a novel model comprised of MultiReader technique and ResNet-GhostVLAD network, which makes the performance on speaker recognition task more stable and excellent. The MultiReader technique and dictionary-based GhostVLAD layer enable our model to filter data at data-level and feature-level respectively, and make it robust for imbalanced datasets which contain noisy and irrelevant information. The proposed model has been trained on the VoxCeleb and AISHELL combined dataset, then tested on AISHELL dataset. Evaluations for different weights of the multiple datasets show that our model outperforms the method of directly mixing training set by obvious margins, which are 10.7% and 44.3% relative improvement.

Similar papers

© 2026 NYSGPT2525 LLC