Adversarial examples in the physical world
2016/07/08 by Alexey Kurakin, Ian Goodfellow, Kurakin, Alexey +3 · 5 voices · 124 citations
Computer Science · Mathematics · #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.CR #cs.CV #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1607.02533
arxiv published 2016/07/08 · arxiv updated 2017/02/11
Abstract
Most existing machine learning classifiers are highly vulnerable to adversarial examples. An adversarial example is a sample of input data which has been modified very slightly in a way that is intended to cause a machine learning classifier to misclassify it. In many cases, these modifications can be so subtle that a human observer does not even notice the modification at all, yet the classifier still makes a mistake. Adversarial examples pose security concerns because they could be used to perform an attack on machine learning systems, even if the adversary has no access to the underlying model. Up to now, all previous work have assumed a threat model in which the adversary can feed data directly into the machine learning classifier. This is not always the case for systems operating in the physical world, for example those which are using signals from cameras and other sensors as an input. This paper shows that even in such physical world scenarios, machine learning systems are vulnerable to adversarial examples. We demonstrate this by feeding adversarial images obtained from cell-phone camera to an ImageNet Inception classifier and measuring the classification accuracy of the system. We find that a large fraction of adversarial examples are classified incorrectly even when perceived through the camera.
Cited by
- GLST: Defending Confidence-Driven V2X Collaborative Perception Against Stealthy Multi-Attacker Feature Injection
- A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection
- Multi-Layer Confidence Scoring for Detection of Out-of-Distribution Samples, Adversarial Attacks, and In-Distribution Misclassifications
- TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
- Towards Transferable Defense Against Malicious Image Edits
- Optimizing the Adversarial Perturbation with a Momentum-based Adaptive Matrix
- Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS
- Calibrating Uncertainty for Zero-Shot Adversarial CLIP
- GradID: Adversarial Detection via Intrinsic Dimensionality of Gradients
- Keep the Lights On, Keep the Lengths in Check: Plug-In Adversarial Detection for Time-Series LLMs in Energy Forecasting
- SPOOF: Simple Pixel Operations for Out-of-Distribution Fooling
- QSTAformer: A Quantum-Enhanced Transformer for Robust Short-Term Voltage Stability Assessment against Adversarial Attacks
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
- Frequency Bias Matters: Diving into Robust and Generalized Deep Image Forgery Detection
- Targeted Manipulation: Slope-Based Attacks on Financial Time-Series Data
- AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
- Robust Physical Adversarial Patches Using Dynamically Optimized Clusters
- A Novel and Practical Universal Adversarial Perturbations against Deep Reinforcement Learning based Intrusion Detection Systems
- Angular Gradient Sign Method: Uncovering Vulnerabilities in Hyperbolic Networks
- Vulnerability-Aware Robust Multimodal Adversarial Training
- Enhancing Adversarial Transferability through Block Stretch and Shrink
- Transferable Dual-Domain Feature Importance Attack against AI-Generated Image Detector
- Cheating Stereo Matching in Full-scale: Physical Adversarial Attack against Binocular Depth Estimation in Autonomous Driving
- MPD-SGR: Robust Spiking Neural Networks with Membrane Potential Distribution-Driven Surrogate Gradient Regularization
- Robust Bidirectional Associative Memory via Regularization Inspired by the Subspace Rotation Algorithm
- On the Trade-Off Between Transparency and Security in Adversarial Machine Learning
- Unsupervised Robust Domain Adaptation: Paradigm, Theory and Algorithm
- A Generative Adversarial Approach to Adversarial Attacks Guided by Contrastive Language-Image Pre-trained Model
- From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge
- Deep learning models are vulnerable, but adversarial examples are even more vulnerable
- Beyond Deceptive Flatness: Dual-Order Solution for Strengthening Adversarial Transferability
- SAAIPAA: Optimizing aspect-angles-invariant physical adversarial attacks on SAR target recognition models
- Diffusion Models are Robust Pretrainers
- Parameter Interpolation Adversarial Training for Robust Image Classification
- Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
- Rethinking Robust Adversarial Concept Erasure in Diffusion Models
- C-LEAD: Contrastive Learning for Enhanced Adversarial Defense
- ANCHOR: Integrating Adversarial Training with Hard-mined Supervised Contrastive Learning for Robust Representation Learning
- Fine-Grained Iterative Adversarial Attacks with Limited Computation Budget
- Test-Time Defense Against Adversarial Attacks via Stochastic Resonance of Latent Ensembles
- Adversarially Robust Quantum Transfer Learning
- A New Type of Adversarial Examples
- Ensuring Robustness in ML-enabled Software Systems: A User Survey
- Investigating Adversarial Robustness against Preprocessing used in Blackbox Face Recognition
- Unveiling the Vulnerability of Graph-LLMs: An Interpretable Multi-Dimensional Adversarial Attack on TAGs
- StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided Illusions
- ROFI: A Deep Learning-Based Ophthalmic Sign-Preserving and Reversible Patient Face Anonymizer
- Mirage Fools the Ear, Mute Hides the Truth: Precise Targeted Adversarial Attacks on Polyphonic Sound Event Detection Systems
- Tight Robustness Certificates and Wasserstein Distributional Attacks for Deep Neural Networks
- A Study of the Removability of Speaker-Adversarial Perturbations
- SynthID-Image: Image watermarking at internet scale
- Robust Driving Control for Autonomous Vehicles: An Intelligent General-sum Constrained Adversarial Reinforcement Learning Approach
- Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-Distillers
- Cyber Resilience Assessment of Unbalanced Distribution System Restoration under Sparse Load Forecasting Attacks
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- On the Adversarial Robustness of Learning-based Conformal Novelty Detection
- Visual CoT Makes VLMs Smarter but More Fragile
- Merge Now, Regret Later: The Hidden Cost of Model Merging is Adversarial Transferability
- Accuracy-Robustness Trade Off via Spiking Neural Network Gradient Sparsity Trail
- Real-World Transferable Adversarial Attack on Face-Recognition Systems
- Targeted perturbations reveal brain-like local coding axes in robustified, but not standard, ANN-based brain models
- Decoding Deception: Understanding Automatic Speech Recognition Vulnerabilities in Evasion and Poisoning Attacks
- Zubov-Net: Adaptive Stability for Neural ODEs Reconciling Accuracy with Robustness
- Position: Human Factors Reshape Adversarial Analysis in Human-AI Decision-Making Systems
- The Use of the Simplex Architecture to Enhance Safety in Deep-Learning-Powered Autonomous Systems
- Understanding and Improving Adversarial Robustness of Neural Probabilistic Circuits
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
- Continual Learning with Vision-Language Models via Semantic-Geometry Preservation
- A Validation Strategy for Deep Learning Models: Evaluating and Enhancing Robustness
- Semantic Representation Attack against Aligned Large Language Models
- Adversarial Examples Are Not Bugs, They Are Superposition
- Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification
- DiffHash: Text-Guided Targeted Attack via Diffusion Models against Deep Hashing Image Retrieval
- Sy-FAR: Symmetry-based Fair Adversarial Robustness
- A Practical Adversarial Attack against Sequence-based Deep Learning Malware Classifiers
- DARD: Dice Adversarial Robustness Distillation against Adversarial Attacks
- A Modern Look at Simplicity Bias in Image Classification Tasks
- MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
- Generating Transferrable Adversarial Examples via Local Mixing and Logits Optimization for Remote Sensing Object Recognition
- SAGE: Sample-Aware Guarding Engine for Robust Intrusion Detection Against Adversarial Attacks
- DisPatch: Disarming Adversarial Patches in Object Detection with Diffusion Models
- A Brain-Inspired Gating Mechanism Unlocks Robust Computation in Spiking Neural Networks
- An Investigation of Visual Foundation Models Robustness
- PromptFlare: Prompt-Generalized Defense via Cross-Attention Decoy in Diffusion-Based Inpainting
- SoK: Understanding the Fundamentals and Implications of Sensor Out-of-band Vulnerabilities
- Get Global Guarantees: On the Probabilistic Nature of Perturbation Robustness
- MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations
- Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization
- Towards Unified Probabilistic Verification and Validation of Vision-Based Autonomy
- Timestep-Compressed Attack on Spiking Neural Networks through Timestep-Level Backpropagation
- DASH: A Meta-Attack Framework for Synthesizing Effective and Stealthy Adversarial Examples
- ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers
- Semantically Guided Adversarial Testing of Vision Models Using Language Models
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- Constrained Black-Box Attacks Against Multi-Agent Reinforcement Learning
- FS-IQA: Certified Feature Smoothing for Robust Image Quality Assessment
- Keep It Real: Challenges in Attacking Compression-Based Adversarial Purification
- Safety of Embodied Navigation: A Survey
- Boosting Adversarial Transferability via Residual Perturbation Attack
- Are Inherently Interpretable Models More Robust? A Study In Music Emotion Recognition
- When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
- The Power of Many: Synergistic Unification of Diverse Augmentations for Efficient Adversarial Robustness
- GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
- Beyond Vulnerabilities: A Survey of Adversarial Attacks as Both Threats and Defenses in Computer Vision Systems
- CP-FREEZER: Latency Attacks against Vehicular Cooperative Perception
- BOOD: Boundary-based Out-Of-Distribution Data Generation
- Adversarial-Guided Diffusion for Multimodal LLM Attacks
- Theoretical Analysis of Relative Errors in Gradient Computations for Adversarial Attacks with CE Loss
- NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples
- Revisiting Physically Realizable Adversarial Object Attack against LiDAR-based Detection: Clarifying Problem Formulation and Experimental Protocols
- The Endless Tuning. An Artificial Intelligence Design To Avoid Human Replacement and Trace Back Responsibilities
- Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics
- Probabilistic Safety Verification for an Autonomous Ground Vehicle: A Situation Coverage Grid Approach
- Kaleidoscopic Background Attack: Disrupting Pose Estimation with Multi-Fold Radial Symmetry Textures
- TRIX- Trading Adversarial Fairness via Mixed Adversarial Training
- VERITAS: Verification and Explanation of Realness in Images for Transparency in AI Systems
- A Rigorous Behavior Assessment of CNNs Using a Data-Domain Sampling Regime
- 3D Gaussian Splatting Driven Multi-View Robust Physical Adversarial Camouflage Generation
- What to Do Next? Memorizing skills from Egocentric Instructional Video
- Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations
- PBCAT: Patch-based composite adversarial training against physically realizable attacks on object detection
- Concept-based Adversarial Attack: a Probabilistic Perspective
- Fragile, Robust, and Antifragile: A Perspective from Parameter Responses in Reinforcement Learning Under Stress
Discussions
- Adversarial examples in the physical world [hn, 3 points, 0 comments]
- Adverserial Examples in the Physical World [hn, 2 points, 0 comments]
- Adversarial Examples in the Real World [hn, 1 points, 0 comments]
- Fooling machine learning image classifier [hn, 1 points, 0 comments]
- Adversarial examples in the physical world [hn, 1 points, 0 comments]
Related