Speaker recognition is one of the biometric technologies to recognize humans, such as identification of fingerprints, DNA, and iris because the characteristics of human speech are unique. The characteristics of speech are dominated by voiced segment. Meanwhile, the unvoiced segments those commonly have random waveform can be seen as noise for the speaker recognition. Therefore, in this paper, a simple procedure to remove unvoiced parts is proposed to improve the accuracy. This procedure is simply implemented using the short time zero-crossing rate (STZCR) to detect and remove the unvoiced parts. It is applied before the Mel Frequency Cepstral Coefficients (MFCC)-based feature extraction and the Gaussian Mixture Model (GMM)-based classification. Evaluation on VoxForge dataset using 4-fold cross-validation show that the proposed procedure is capable of improving the accuracy from 99.94% to 100%.
Paper
Full text
Removing Unvoiced Segment to Improve Text Independent Speaker Recognition
Semantic Scholar · Computer Science · 2019
Abstract
Speaker recognition is one of the biometric technologies to recognize humans, such as identification of fingerprints, DNA, and iris because the characteristics of human speech are unique. The characteristics of speech are dominated by voiced segment. Meanwhile, the unvoiced segments those commonly have random waveform can be seen as noise for the speaker recognition. Therefore, in this paper, a simple procedure to remove unvoiced parts is proposed to improve the accuracy. This procedure is simply implemented using the short time zero-crossing rate (STZCR) to detect and remove the unvoiced parts. It is applied before the Mel Frequency Cepstral Coefficients (MFCC)-based feature extraction and the Gaussian Mixture Model (GMM)-based classification. Evaluation on VoxForge dataset using 4-fold cross-validation show that the proposed procedure is capable of improving the accuracy from 99.94% to 100%.