Representation Learning to Classify and Detect Adversarial Attacks\n against Speaker and Speech Recognition Systems

Adversarial attacks have become a major threat for machine learning\napplications. There is a growing interest in studying these attacks in the\naudio domain, e.g, speech and speaker recognition; and find defenses against\nthem. In this work, we focus on using representation learning to\nclassify/detect attacks w.r.t. the attack algorithm, threat model or\nsignal-to-adversarial-noise ratio. We found that common attacks in the\nliterature can be classified with accuracies as high as 90%. Also,\nrepresentations trained to classify attacks against speaker identification can\nbe used also to classify attacks against speaker verification and speech\nrecognition. We also tested an attack verification task, where we need to\ndecide whether two speech utterances contain the same attack. We observed that\nour models did not generalize well to attack algorithms not included in the\nattack representation model training. Motivated by this, we evaluated an\nunknown attack detection task. We were able to detect unknown attacks with\nequal error rates of about 19%, which is promising.\n

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC