SoK: The Faults in our ASRs: An Overview of Attacks against Automatic Speech Recognition and Speaker Identification Systems

Speech and speaker recognition systems are employed in a variety of\napplications, from personal assistants to telephony surveillance and biometric\nauthentication. The wide deployment of these systems has been made possible by\nthe improved accuracy in neural networks. Like other systems based on neural\nnetworks, recent research has demonstrated that speech and speaker recognition\nsystems are vulnerable to attacks using manipulated inputs. However, as we\ndemonstrate in this paper, the end-to-end architecture of speech and speaker\nsystems and the nature of their inputs make attacks and defenses against them\nsubstantially different than those in the image space. We demonstrate this\nfirst by systematizing existing research in this space and providing a taxonomy\nthrough which the community can evaluate future work. We then demonstrate\nexperimentally that attacks against these models almost universally fail to\ntransfer. In so doing, we argue that substantial additional work is required to\nprovide adequate mitigations in this space.\n

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC