Generating Out of Distribution Adversarial Attack using Latent Space Poisoning

Traditional adversarial attacks rely upon the perturbations generated by gradients from the network which are generally safeguarded by gradient guided search to provide an adversarial counterpart to the network. In this letter, we propose a novel framework to generate adversarial examples where the actual image is not corrupted rather its latent space representation is utilized to tamper the inherent structure of the image while maintaining the perceptual quality intact and to act as legitimate data samples. As opposed to gradient-based attacks, the latent space poisoning exploits the inclination of classifiers to model the independent and identical distribution of the training dataset and tricks it by producing out of distribution samples. We train a disentangled variational autoencoder (<inline-formula><tex-math notation="LaTeX">$\beta$</tex-math></inline-formula>-VAE) to model the data in latent space and then we add noise perturbations using a class-conditioned distribution function to the latent space under the constraint that it is misclassified to the target label. Our empirical results on MNIST, SVHN, and CelebA dataset validate that the generated adversarial examples can easily fool robust <inline-formula><tex-math notation="LaTeX">$l_0$</tex-math></inline-formula>, <inline-formula><tex-math notation="LaTeX">$l_2$</tex-math></inline-formula>, <inline-formula><tex-math notation="LaTeX">$l_{\infty }$</tex-math></inline-formula> norm classifiers designed using provably robust defense mechanisms. The source code is made publicly available at <uri>https://github.com/Ujjwal-9/latent-space-poisoning</uri>

Paper

References (40)

Scroll for more · 28 remaining

Similar papers

© 2026 NYSGPT2525 LLC