We make inroads into understanding the robustness of Variational Autoencoders\n(VAEs) to adversarial attacks and other input perturbations. While previous\nwork has developed algorithmic approaches to attacking and defending VAEs,\nthere remains a lack of formalization for what it means for a VAE to be robust.\nTo address this, we develop a novel criterion for robustness in probabilistic\nmodels: $r$-robustness. We then use this to construct the first theoretical\nresults for the robustness of VAEs, deriving margins in the input space for\nwhich we can provide guarantees about the resulting reconstruction. Informally,\nwe are able to define a region within which any perturbation will produce a\nreconstruction that is similar to the original reconstruction. To support our\nanalysis, we show that VAEs trained using disentangling methods not only score\nwell under our robustness metrics, but that the reasons for this can be\ninterpreted through our theoretical results.\n
Paper
References (35)
Scroll for more · 23 remaining