Gradient-based adversarial attacks on categorical sequence models via traversing an embedded world

Deep learning models suffer from a phenomenon called adversarial attacks: we\ncan apply minor changes to the model input to fool a classifier for a\nparticular example. The literature mostly considers adversarial attacks on\nmodels with images and other structured inputs. However, the adversarial\nattacks for categorical sequences can also be harmful. Successful attacks for\ninputs in the form of categorical sequences should address the following\nchallenges: (1) non-differentiability of the target function, (2) constraints\non transformations of initial sequences, and (3) diversity of possible\nproblems. We handle these challenges using two black-box adversarial attacks.\nThe first approach adopts a Monte-Carlo method and allows usage in any\nscenario, the second approach uses a continuous relaxation of models and target\nmetrics, and thus allows usage of state-of-the-art methods for adversarial\nattacks with little additional effort. Results for money transactions, medical\nfraud, and NLP datasets suggest that proposed methods generate reasonable\nadversarial sequences that are close to original ones but fool machine learning\nmodels.\n

Paper

References (42)

Scroll for more · 30 remaining

Similar papers

© 2026 NYSGPT2525 LLC