Improving on-device speaker verification using federated learning with privacy

Information on speaker characteristics can be useful as side information in\nimproving speaker recognition accuracy. However, such information is often\nprivate. This paper investigates how privacy-preserving learning can improve a\nspeaker verification system, by enabling the use of privacy-sensitive speaker\ndata to train an auxiliary classification model that predicts vocal\ncharacteristics of speakers. In particular, this paper explores the utility\nachieved by approaches which combine different federated learning and\ndifferential privacy mechanisms. These approaches make it possible to train a\ncentral model while protecting user privacy, with users' data remaining on\ntheir devices. Furthermore, they make learning on a large population of\nspeakers possible, ensuring good coverage of speaker characteristics when\ntraining a model. The auxiliary model described here uses features extracted\nfrom phrases which trigger a speaker verification system. From these features,\nthe model predicts speaker characteristic labels considered useful as side\ninformation. The knowledge of the auxiliary model is distilled into a speaker\nverification system using multi-task learning, with the side information labels\npredicted by this auxiliary model being the additional task. This approach\nresults in a 6% relative improvement in equal error rate over a baseline\nsystem.\n

Paper

References (32)

Scroll for more · 20 remaining

Similar papers

© 2026 NYSGPT2525 LLC