Deep neural network approaches to speaker verification have proven\nsuccessful, but typical computational requirements of State-Of-The-Art (SOTA)\nsystems make them unsuited for embedded applications. In this work, we present\na two-stage model architecture orders of magnitude smaller than common\nsolutions (237.5K learning parameters, 11.5MFLOPS) reaching a competitive\nresult of 3.31% Equal Error Rate (EER) on the well established VoxCeleb1\nverification test set. We demonstrate the possibility of running our solution\non small devices typical of IoT systems such as the Raspberry Pi 3B with a\nlatency smaller than 200ms on a 5s long utterance. Additionally, we evaluate\nour model on the acoustically challenging VOiCES corpus. We report a limited\nincrease in EER of 2.6 percentage points with respect to the best scoring model\nof the 2019 VOiCES from a Distance Challenge, against a reduction of 25.6 times\nin the number of learning parameters.\n
Paper
References (27)
Scroll for more · 15 remaining