2018/11/20 by Qian Huang, Huang, Qian, Zeqi Gu +13 · 1 citation
Computer Science · Materials Science · Mathematics · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Adversarial system #Artificial intelligence #Artificial neural network #Black box #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer security #Deep neural networks #Exploit #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Nuclear Materials and Properties #Overfitting #Representation (politics) #Source model #Theoretical computer science #Transfer of learning #Transferability #cs.CV #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1811.08458
Preprint
arxiv created 2018/11/20 · openalex publication_date 2018/11/20 · arxiv updated 2018/11/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Neural networks are vulnerable to adversarial examples, malicious inputs crafted to fool trained models. Adversarial examples often exhibit black-box transfer, meaning that adversarial examples for one model can fool another model. However, adversarial examples may be overfit to exploit the particular architecture and feature representation of a source model, resulting in sub-optimal black-box transfer attacks to other target models. This leads us to introduce the Intermediate Level Attack (ILA), which attempts to fine-tune an existing adversarial example for greater black-box transferability by increasing its perturbation on a pre-specified layer of the source model. We show that our method can effectively achieve this goal and that we can decide a nearly-optimal layer of the source model to perturb without any knowledge of the target models.