Towards an Efficient Voice Identification Using Wav2Vec2.0 and HuBERT Based on the Quran Reciters Dataset

Current authentication and trusted systems depend on classical and biometric\nmethods to recognize or authorize users. Such methods include audio speech\nrecognitions, eye, and finger signatures. Recent tools utilize deep learning\nand transformers to achieve better results. In this paper, we develop a deep\nlearning constructed model for Arabic speakers identification by using\nWav2Vec2.0 and HuBERT audio representation learning tools. The end-to-end\nWav2Vec2.0 paradigm acquires contextualized speech representations learnings by\nrandomly masking a set of feature vectors, and then applies a transformer\nneural network. We employ an MLP classifier that is able to differentiate\nbetween invariant labeled classes. We show several experimental results that\nsafeguard the high accuracy of the proposed model. The experiments ensure that\nan arbitrary wave signal for a certain speaker can be identified with 98% and\n97.1% accuracies in the cases of Wav2Vec2.0 and HuBERT, respectively.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC