End to end speaker embedding systems have shown promising performance on speaker recognition tasks. However, speaker recognition system often degrade obviously when there are different scale and discrepant distribution of training datasets. In this paper, we proposed a novel model comprised of MultiReader technique and ResNet-GhostVLAD network, which makes the performance on speaker recognition task more stable and excellent. The MultiReader technique and dictionary-based GhostVLAD layer enable our model to filter data at data-level and feature-level respectively, and make it robust for imbalanced datasets which contain noisy and irrelevant information. The proposed model has been trained on the VoxCeleb and AISHELL combined dataset, then tested on AISHELL dataset. Evaluations for different weights of the multiple datasets show that our model outperforms the method of directly mixing training set by obvious margins, which are 10.7% and 44.3% relative improvement.
Paper
Full text
Multi-level Data Filter for Speaker Recognition on Imbalanced Datasets
Semantic Scholar · Computer Science · 2019
Abstract
End to end speaker embedding systems have shown promising performance on speaker recognition tasks. However, speaker recognition system often degrade obviously when there are different scale and discrepant distribution of training datasets. In this paper, we proposed a novel model comprised of MultiReader technique and ResNet-GhostVLAD network, which makes the performance on speaker recognition task more stable and excellent. The MultiReader technique and dictionary-based GhostVLAD layer enable our model to filter data at data-level and feature-level respectively, and make it robust for imbalanced datasets which contain noisy and irrelevant information. The proposed model has been trained on the VoxCeleb and AISHELL combined dataset, then tested on AISHELL dataset. Evaluations for different weights of the multiple datasets show that our model outperforms the method of directly mixing training set by obvious margins, which are 10.7% and 44.3% relative improvement.