Shortcut learning in deep neural networks
2020/04/16 by Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis +5 · 2 voices · 162 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Machine Learning and Data Classification #cs.AI #cs.CV #cs.LG #q-bio.NC
paper · pdf · doi:10.1038/s42256-020-00257-z
openalex publication_date 2020/11/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
Deep learning has triggered the current rise of artificial intelligence and is the workhorse of today's machine intelligence. Numerous success stories have rapidly spread all over science, industry and society, but its limitations have only recently come into focus. In this perspective we seek to distill how many of deep learning's problems can be seen as different symptoms of the same underlying problem: shortcut learning. Shortcuts are decision rules that perform well on standard benchmarks but fail to transfer to more challenging testing conditions, such as real-world scenarios. Related issues are known in Comparative Psychology, Education and Linguistics, suggesting that shortcut learning may be a common characteristic of learning systems, biological and artificial alike. Based on these observations, we develop a set of recommendations for model interpretation and benchmarking, highlighting recent advances in machine learning to improve robustness and transferability from the lab to real-world applications.
Citations
Cited by
- A Compression Perspective on Simplicity Bias
- Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification
- G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement
- Training Language Models via Neural Cellular Automata
- DLMUSE: Robust Brain Segmentation in Seconds Using Deep Learning
- A Pipeline for Automated Quality Control of Chest Radiographs
- A Framework for Autonomous AI-Driven Drug Discovery
- Robustness and trustworthiness in AI: a no-go result from formal epistemology
- Contrastive learning enhances fairness in pathology artificial intelligence systems
- Deep problems with neural network models of human vision
- Better Data and Smarter AI: Automated Quality Control for Chest Radiographs
- Compositionality and Sentence Meaning: Comparing Semantic Parsing and Transformers on a Challenging Sentence Similarity Dataset
- Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- Usefulness of interpretability methods to explain deep learning based plant stress phenotyping
- An Invitation to Deep Reinforcement Learning
- Environment Inference for Invariant Learning
- The Scaling Properties of Implicit Deductive Reasoning in Transformers
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- Natural Adversarial Examples
- Process Matters more than Output for Distinguishing Humans from Machines
- Fighting Copycat Agents in Behavioral Cloning from Observation Histories
- Symmetry and Generalisation in Neural Approximations of Renormalisation Transformations
- Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
- An Empirical Study of Foundation Models for Variability-Induced Compilation Errors in Configurable C Code
- White Paper Assistance: A Step Forward Beyond the Shortcut Learning
- Debiasing Methods in Natural Language Understanding Make Bias More\n Accessible
- HistoLens: An Interactive XAI Toolkit for Verifying and Mitigating Flaws in Vision-Language Models for Histopathology
- Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders
- Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
- Preventing Shortcuts in Adapter Training via Providing the Shortcuts
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
- ShapeX: Shapelet-Driven Post Hoc Explanations for Time Series Classification Models
- Vision-language models learn the geometry of human perceptual space
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- Overinterpretation reveals image classification model pathologies
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining
- Beyond Frequency: Scoring-Driven Debiasing for Object Detection via Blueprint-Prompted Image Synthesis
- How human–AI feedback loops alter human perceptual, emotional and social judgements
- The risk of shortcutting in deep learning algorithms for medical imaging research
- The debate over understanding in AI’s large language models
- Improving statistical precision in Monte Carlo samples with negative weights via reweighting and uncertainty quantification
- Scaling Artificial Intelligence for Multi-Tumor Early Detection with More Reports, Fewer Masks
- LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching
- ESI: Epistemic Uncertainty Quantification via Semantic-preserving Intervention for Large Language Models
- Novel Class Discovery for Point Cloud Segmentation via Joint Learning of Causal Representation and Reasoning
- Hybrid Explanation-Guided Learning for Transformer-Based Chest X-Ray Diagnosis
- FACE: Faithful Automatic Concept Extraction
- LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
- Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
- Gradient-based Model Shortcut Detection for Time Series Classification
- Explainable Human-in-the-Loop Segmentation via Critic Feedback Signals
- Towards Neurocognitive-Inspired Intelligence: From AI's Structural Mimicry to Human-Like Functional Cognition
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
- Revisiting Mixout: An Overlooked Path to Robust Finetuning
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
- Learning What Matters: Steering Diffusion via Spectrally Anisotropic Forward Noise
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance
- Robust Context-Aware Object Recognition
- Diagnosing Shortcut-Induced Rigidity in Continual Learning: The Einstellung Rigidity Index (ERI)
- Exploring System 1 and 2 communication for latent reasoning in LLMs
- Cracking the code of adaptive immunity: The role of computational tools
- Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
- Causally Guided Gaussian Perturbations for Out-Of-Distribution Generalization in Medical Imaging
- Successful Misunderstandings: Learning to Coordinate Without Being Understood
- Fidelity-Aware Data Composition for Robust Robot Generalization
- Tasting the cake: evaluating self-supervised generalization on\n out-of-distribution multimodal MRI data
- GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
- RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
- Probabilistic Numeric Convolutional Neural Networks
- ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech
- Interpretable deep learning: interpretation, interpretability, trustworthiness, and beyond
- Causally-Enhanced Reinforcement Policy Optimization
- When Can AI Models Explain Learning? Validity Criteria for AI as Cognitive Models in Education
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing
- Review of Hallucination Understanding in Large Language and Vision Models
- Learning to Look: Cognitive Attention Alignment with Vision-Language Models
- Exploring the Limits of Out-of-Distribution Detection
- Understanding and Improving Adversarial Robustness of Neural Probabilistic Circuits
- Artificial Intelligence’s new clothes? A system technology perspective
- Scaling Vision-Language Models Is Not Enough to Mitigate Bias
- What's in a Name? Morphological Shortcuts by LLMs in Pharmacology
- What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
- Modeling Decisions in Blockchain Analytics: A Leakage-Aware Evaluation of Tree-Based vs. Sequential Models
- Context-Informed Ship Trajectory Prediction via Conditional Attention
- Compression-Based Behavioral Similarity for Open-World Sybil Discovery on Ethereum
- Shortcut to Nowhere: Demystifying Deep Spurious Regression
- Selective Forgetting of Deep Networks at a Finer Level than Samples
- A Guide to Metabolic Network Modeling for Plant Biology
- Dimensions underlying the representational alignment of deep neural networks with humans
- Saliency is a Possible Red Herring When Diagnosing Poor Generalization
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- AI Error Difficulty Modulates the Effectiveness of Explainability in Decision Support Systems
- Effect of Demographic Bias on Skin Lesion Classification
- BB-EIT: A Generalized Prediction Model for Protein Adsorption on Polymer Brushes Using Augmented Chemical Embeddings
- Vision-language model-based semantic-guided imaging biomarker for lung nodule malignancy prediction
- Debugging Concept Bottleneck Models through Removal and Retraining
- Probabilistic Runtime Verification, Evaluation and Risk Assessment of Visual Deep Learning Systems
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- Implicit Regularization via Neural Feature Alignment
- Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
- End2Race: Efficient End-to-End Imitation Learning for Real-Time F1Tenth Racing
- From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?
- Adapting Machine Learning Diagnostic Models to New Populations Using a Small Amount of Data: Results from Clinical Neuroscience
- Data Scaling Laws for Radiology Foundation Models
- Causal-Symbolic Meta-Learning (CSML): Inducing Causal World Models for Few-Shot Generalization
- Data augmentation and image understanding
- Disentangling Content from Style to Overcome Shortcut Learning: A Hybrid Generative-Discriminative Learning Framework
- Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble
- Invisible Attributes, Visible Biases: Exploring Demographic Shortcuts in MRI-based Alzheimer's Disease Classification
- Kriging prior Regression: A Case for Kriging-Based Spatial Features with TabPFN in Soil Mapping
- Bias-Aware Machine Unlearning: Towards Fairer Vision Models via Controllable Forgetting
- ACE and Diverse Generalization via Selective Disagreement
- PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology
- A guide to machine learning for biologists
- The Advantage of Fine-Grained Training
- COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
- Ecologically Valid Benchmarking and Adaptive Attention: Scalable Marine Bioacoustic Monitoring
- An Approach to Grounding AI Model Evaluations in Human-derived Criteria
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- Differentiable Entropy Regularization: A Complexity-Aware Approach for Neural Optimization
- Multi Attribute Bias Mitigation via Representation Learning
- Rashomon in the Streets: Explanation Ambiguity in Scene Understanding
- WILDS: A Benchmark of in-the-Wild Distribution Shifts
- Partial success in closing the gap between human and machine vision
- PreferenceNet: Encoding Human Preferences in Auction Design with Deep Learning
- Pose Discrepancy Spatial Transformer Based Feature Disentangling for Partial Aspect Angles SAR Target Recognition
- On Disentangled Representations Learned From Correlated Data
- NM-Hebb: Coupling Local Hebbian Plasticity with Metric Learning for More Accurate and Interpretable CNNs
- Gradient Rectification for Robust Calibration under Distribution Shift
- Quantum Entanglement as Super-Confounding: From Bell's Theorem to Robust Machine Learning
- Self-Supervised Learning with Data Augmentations Provably Isolates\n Content from Style
- From Prediction to Simulation: AlphaFold 3 as a Differentiable Framework for Structural Biology
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
- Wormhole Dynamics in Deep Neural Networks
- Mitigating Easy Option Bias in Multiple-Choice Question Answering
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- SimAQ: Mitigating Experimental Artifacts in Soft X-Ray Tomography using Simulated Acquisitions
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
- Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm Detection
- Improving ARDS Diagnosis Through Context-Aware Concept Bottleneck Models
- How benign is benign overfitting?
- Position: Causal Machine Learning Requires Rigorous Synthetic Experiments for Broader Adoption
- Mitigating Biases in Surgical Operating Rooms with Geometry
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- Class Unbiasing for Generalization in Medical Diagnosis
- From Explainable to Explained AI: Ideas for Falsifying and Quantifying Explanations
- Simple data balancing achieves competitive worst-group-accuracy
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- Simulating Human-Like Learning Dynamics with LLM-Empowered Agents
- Learning Robust Intervention Representations with Delta Embeddings
- Multilingual Source Tracing of Speech Deepfakes: A First Benchmark
- Versatile yet Efficient Network Traffic Analysis: Offloading Network Foundation Model to SmartNIC
- Evading Data Provenance in Deep Neural Networks
- Foundations of Interpretable Models
- Forgetting of task-specific knowledge in model merging-based continual learning
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- Mitigating Spurious Correlations in Weakly Supervised Semantic Segmentation via Cross-architecture Consistency Regularization
- Self-supervision of Feature Transformation for Further Improving Supervised Learning
Discussions
Related