2018/05/30 by Dimitris Tsipras, Tsipras, Dimitris, Shibani Santurkar +8 · 168 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE) #cs.CV #cs.LG #cs.NE #stat.ML
paper · pdf · doi:10.48550/arxiv.1805.12152
ICLR'19
openalex publication_date 2018/05/30 · arxiv created 2019/09/09 · arxiv updated 2019/09/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We show that there may exist an inherent tension between the goal of adversarial robustness and that of standard generalization. Specifically, training robust models may not only be more resource-consuming, but also lead to a reduction of standard accuracy. We demonstrate that this trade-off between the standard accuracy of a model and its robustness to adversarial perturbations provably exists in a fairly simple and natural setting. These findings also corroborate a similar phenomenon observed empirically in more complex settings. Further, we argue that this phenomenon is a consequence of robust classifiers learning fundamentally different feature representations than standard classifiers. These differences, in particular, seem to result in unexpected benefits: the representations learned by robust models tend to align better with salient data characteristics and human perception.