Deep Neural Networks (DNNs) have been widely used in speech processing and show great performance on a range of tasks, such as speech recognition, machine translation, and speaker verification. In this paper we propose a new type of DNN model for text-dependent speaker verification. The frame-level features, being extracted by DNN, are usually equally weighted and aggregated (or averaged) to compute an utterance-level speaker representation. We combine Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) to extract frame-level features, then provide every frame with different weights produced by attention mechanism, so that utterance-level speaker representation can be generated by weighted averaging frame-level features. We explore different generating methods on the attention weights. Besides, attention mechanism can also be used to align temporal information between enrollment and evaluation utterance. Triplet loss function is used to optimize our models, requiring inputs group in triplet style. Ultimately, results of experiment on RSR2015 database show that our attention-based model outperforms various baseline models.
Paper
Full text
Attentional triplet neural networks for text-dependent speaker verification
OpenAlex · Speech Recognition and Synthesis · 2020
Abstract
Deep Neural Networks (DNNs) have been widely used in speech processing and show great performance on a range of tasks, such as speech recognition, machine translation, and speaker verification. In this paper we propose a new type of DNN model for text-dependent speaker verification. The frame-level features, being extracted by DNN, are usually equally weighted and aggregated (or averaged) to compute an utterance-level speaker representation. We combine Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) to extract frame-level features, then provide every frame with different weights produced by attention mechanism, so that utterance-level speaker representation can be generated by weighted averaging frame-level features. We explore different generating methods on the attention weights. Besides, attention mechanism can also be used to align temporal information between enrollment and evaluation utterance. Triplet loss function is used to optimize our models, requiring inputs group in triplet style. Ultimately, results of experiment on RSR2015 database show that our attention-based model outperforms various baseline models.
References (21)
Scroll for more · 9 remaining