Adversarial Examples Are Not Bugs, They Are Features
2019/05/06 by Andrew Ilyas, Ilyas, Andrew, Shibani Santurkar +10 · 6 voices · 117 citations
Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #cs.CR #cs.CV #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1905.02175
openalex publication_date 2019/05/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Adversarial examples have attracted significant attention in machine learning, but the reasons for their existence and pervasiveness remain unclear. We demonstrate that adversarial examples can be directly attributed to the presence of non-robust features: features derived from patterns in the data distribution that are highly predictive, yet brittle and incomprehensible to humans. After capturing these features within a theoretical framework, we establish their widespread existence in standard datasets. Finally, we present a simple setting where we can rigorously tie the phenomena we observe in practice to a misalignment between the (human-specified) notion of robustness and the inherent geometry of the data.
Citations
Cited by
- Detectors Learn the Wrong Thing: Shortcut-Resistant Adversarial Training Against Physically Realizable Attacks
- Beyond the reducing valve: towards a computational neurophenomenology of altered states via deep neural networks
- Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
- Geometric 2D Scene Graph Generation
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- Large Language Models as a (Bad) Security Norm in the Context of Regulation and Compliance
- Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
- Sample-wise Adaptive Weighting for Transfer Consistency in Adversarial Distillation
- AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing
- Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- Robust Physical Adversarial Patches Using Dynamically Optimized Clusters
- ATAC: Augmentation-Based Test-Time Adversarial Correction for CLIP
- Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
- Learning Fourier shapes to probe the geometric world of deep neural networks
- C-LEAD: Contrastive Learning for Enhanced Adversarial Defense
- Sparse Model Inversion: Efficient Inversion of Vision Transformers for Data-Free Applications
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
- Bilevel Models for Adversarial Learning and A Case Study
- Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
- Dual-Domain Constraints: Designing Covert and Efficient Adversarial Examples for Secure Communication
- A Versatile Framework for Designing Group-Sparse Adversarial Attacks
- Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
- FrameShield: Adversarially Robust Video Anomaly Detection
- Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization
- Revisiting the Relation Between Robustness and Universality
- The Black Tuesday Attack: how to crash the stock market with adversarial examples to financial forecasting models
- When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
- Readout Representation: Redefining Neural Codes by Input Recovery
- Adversarial Attacks Leverage Interference Between Features in Superposition
- The Easy Path to Robustness: Coreset Selection using Sample Hardness
- Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- SPATA: Systematic Pattern Analysis for Detailed and Transparent Data Cards
- Targeted perturbations reveal brain-like local coding axes in robustified, but not standard, ANN-based brain models
- Sparse Representations Improve Adversarial Robustness of Neural Network Classifiers
- What Does Your Benchmark Really Measure? A Framework for Robust Inference of AI Capabilities
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Robustness Feature Adapter for Efficient Adversarial Training
- Adversarial Examples Are Not Bugs, They Are Superposition
- Data coarse graining can improve model performance
- A Modern Look at Simplicity Bias in Image Classification Tasks
- A Discrepancy-Based Perspective on Dataset Condensation
- Beyond Output Faithfulness: Learning Attributions that Preserve Computational Pathways
- Does simple trump complex? Comparing strategies for adversarial robustness in DNNs
- TriQDef: Disrupting Semantic and Gradient Alignment to Prevent Adversarial Patch Transferability in Quantized Neural Networks
- Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
- EFU: Enforcing Federated Unlearning via Functional Encryption
- Training and Inference within 1 Second -- Tackle Cross-Sensor Degradation of Real-World Pansharpening with Efficient Residual Feature Tailoring
- Semantics Preserving Adversarial Learning
- Learning Dynamics of Attention: Human Prior for Interpretable Machine Reasoning
- The Origins and Prevalence of Texture Bias in Convolutional Neural Networks
- SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures
- ETA: Energy-based Test-time Adaptation for Depth Completion
- Keep It Real: Challenges in Attacking Compression-Based Adversarial Purification
- Data Driven Insights into Composition Property Relationships in FCC High Entropy Alloys
- Are Inherently Interpretable Models More Robust? A Study In Music Emotion Recognition
- The Power of Many: Synergistic Unification of Diverse Augmentations for Efficient Adversarial Robustness
- Learning to Disentangle Robust and Vulnerable Features for Adversarial Detection
- LeakyCLIP: Extracting Training Data from CLIP
- AUV-Fusion: Cross-Modal Adversarial Fusion of User Interactions and Visual Perturbations Against VARS
- On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- Radio Adversarial Attacks on EMG-based Gesture Recognition Networks
- Adversarial Examples and the Deeper Riddle of Induction: The Need for a\n Theory of Artifacts in Deep Learning
- Towards a Robust Deep Neural Network in Texts: A Survey
- Beneficial Perturbations Network for Defending Adversarial Examples
- Adversarial Attacks and Defenses: An Interpretation Perspective
- The Endless Tuning. An Artificial Intelligence Design To Avoid Human Replacement and Trace Back Responsibilities
- Training Meta-Surrogate Model for Transferable Adversarial Attack
- Internal Wasserstein Distance for Adversarial Attack and Defense
- And/or trade-off in artificial neurons: impact on adversarial robustness
- Invisible Backdoor Attacks on Deep Neural Networks via Steganography and Regularization
- On the Interaction of Compressibility and Adversarial Robustness
- Understanding (Non-)Robust Feature Disentanglement and the Relationship Between Low- and High-Dimensional Adversarial Attacks
- Emergence of Quantised Representations Isolated to Anisotropic Functions
- Adversarial machine learning :
- Robustifying ℓ_∞ Adversarial Training to the Union of Perturbation Models
- Universal adversarial examples in speech command classification
- The Probabilistic Fault Tolerance of Neural Networks in the Continuous\n Limit
- An Empirical Evaluation of Adversarial Robustness under Transfer Learning
- Robust Local Features for Improving the Generalization of Adversarial Training
- Attack to Fool and Explain Deep Networks
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
- Holmes: Towards Effective and Harmless Model Ownership Verification to Personalized Large Vision Models via Decoupling Common Features
- SpaNN: Detecting Multiple Adversarial Patches on CNNs by Spanning Saliency Thresholds
- PASS: Private Attributes Protection with Stochastic Data Substitution
- Classification and Adversarial examples in an Overparameterized Linear Model: A Signal Processing Perspective
- If MaxEnt RL is the Answer, What is the Question?
- Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
- Intriguing properties of adversarial training at scale
- SDN-Based False Data Detection With Its Mitigation and Machine Learning Robustness for In-Vehicle Networks
- Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction
- Identifying and Understanding Cross-Class Features in Adversarial Training
- Auditing and Debugging Deep Learning Models via Decision Boundaries: Individual-level and Group-level Analysis
- Calibrated neighborhood aware confidence measure for deep metric learning
- Adversarial Machine Learning in Wireless Communications using RF Data: A Review
- Artificial Intelligence Strategies for National Security and Safety Standards
- A Self-supervised Approach for Adversarial Robustness
- BIRD: Behavior Induction via Representation-structure Distillation
- Are classical deep neural networks weakly adversarially robust?
- What is Adversarial Training for Diffusion Models?
- Your Classifier Can Do More: Towards Bridging the Gaps in Classification, Robustness, and Generation
- The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
- Are Time-Series Foundation Models Deployment-Ready? A Systematic Study of Adversarial Robustness Across Domains
- In defence of mathematical content
- Adversarial Attack and Defense in Deep Ranking
- Ignition Phase : Standard Training for Fast Adversarial Robustness
- To Transfer or Not to Transfer: Misclassification Attacks Against\n Transfer Learned Text Classifiers
- EdgeAgentX: A Novel Framework for Agentic AI at the Edge in Military Communication Networks
- Training on Plausible Counterfactuals Removes Spurious Correlations
- Adversarially Pretrained Transformers may be Universally Robust In-Context Learners
- ShortcutProbe: Probing Prediction Shortcuts for Learning Robust Models
- Use as Many Surrogates as You Want: Selective Ensemble Attack to Unleash Transferability without Sacrificing Resource Efficiency
- Now You See It, Now You Dont: Adversarial Vulnerabilities in Computational Pathology
- When Explainability Meets Adversarial Learning: Detecting Adversarial Examples using SHAP Signatures
- How Can We Characterize Human Generalization and Distinguish It From Generalization in Machines?
- Gödel's Sentence Is An Adversarial Example But Unsolvable
- Rearchitecting Classification Frameworks For Increased Robustness
Discussions
- Adversarial Examples Are Not Bugs, They Are Features [hn, 76 points, 28 comments]
- This is reminding me of this paper about adversarial example in image classification, and how they’re tied to unintuitive features arxiv.org/abs/1905.02175 [bsky, 2 points, 1 comments]
- Cool. I haven't read that. I was thinking of arxiv.org/pdf/1905.02175 [bsky, 1 points, 1 comments]
- ほーん 後で読む arxiv.org/abs/1905.02175 [bsky, 1 points, 0 comments]
- The paper argues that adversarial examples are due to fragile, non-robust features in machine learning, showing their prevalence in datasets and addressing the disconnect between human notions of robu [bsky, 0 points, 0 comments]
- みてた: arxiv.org/abs/1905.02175 [bsky, 0 points, 0 comments]
Related