2020/04/22 by Filipe Condessa, J. Zico Kolter, Zico Kolter +2 · 1 citation
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Adversarial system #Adversary #Artificial intelligence #Artificial neural network #Computer science #Deep neural networks #Digital Media Forensic Detection #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Generative grammar #Generative model #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Mathematics #Robustness (evolution) #Upper and lower bounds #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2004.10608
published in arXiv (Cornell University) (Cornell University)
arxiv created 2020/04/22 · openalex publication_date 2020/04/22 · arxiv updated 2020/04/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Recent work in adversarial attacks has developed provably robust methods for training deep neural network classifiers. However, although they are often mentioned in the context of robustness, deep generative models themselves have received relatively little attention in terms of formally analyzing their robustness properties. In this paper, we propose a method for training provably robust generative models, specifically a provably robust version of the variational auto-encoder (VAE). To do so, we first formally define a (certifiably) robust lower bound on the variational lower bound of the likelihood, and then show how this bound can be optimized during training to produce a robust VAE. We evaluate the method on simple examples, and show that it is able to produce generative models that are substantially more robust to adversarial attacks (i.e., an adversary trying to perturb inputs so as to drastically lower their likelihood under the model).