A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks
2018/07/10 by Lee, Kimin, Kibok Lee, Honglak Lee +4 · 192 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.1807.03888
Abstract
Detecting test samples drawn sufficiently far away from the training distribution statistically or adversarially is a fundamental requirement for deploying a good classifier in many real-world machine learning applications. However, deep neural networks with the softmax classifier are known to produce highly overconfident posterior distributions even for such abnormal samples. In this paper, we propose a simple yet effective method for detecting any abnormal samples, which is applicable to any pre-trained softmax neural classifier. We obtain the class conditional Gaussian distributions with respect to (low- and upper-level) features of the deep models under Gaussian discriminant analysis, which result in a confidence score based on the Mahalanobis distance. While most prior methods have been evaluated for detecting either out-of-distribution or adversarial samples, but not both, the proposed method achieves the state-of-the-art performances for both cases in our experiments. Moreover, we found that our proposed method is more robust in harsh cases, e.g., when the training dataset has noisy labels or small number of samples. Finally, we show that the proposed method enjoys broader usage by applying it to class-incremental learning: whenever out-of-distribution samples are detected, our classification rule can incorporate new classes well without further training deep models.
Cited by
- A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection
- Multi-Layer Confidence Scoring for Detection of Out-of-Distribution Samples, Adversarial Attacks, and In-Distribution Misclassifications
- MAD-OOD: A Deep Learning Cluster-Driven Framework for an Out-of-Distribution Malware Detection and Classification
- VICTOR: Dataset Copyright Auditing in Video Recognition Systems
- Out-of-Distribution Detection for Continual Learning: Design Principles and Benchmarking
- General OOD Detection via Model-aware and Subspace-aware Variable Priority
- Predictive Sample Assignment for Semantically Coherent Out-of-Distribution Detection
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification
- Contrast transfer functions help quantify neural network out-of-distribution generalization in HRTEM
- Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts
- HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability
- Knowing when to trust machine-learned interatomic potentials
- Domain Feature Collapse: Implications for Out-of-Distribution Detection and Solutions
- Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs
- Studying Various Activation Functions and Non-IID Data for Machine Learning Model Robustness
- Credal Graph Neural Networks
- TIE: A Training-Inversion-Exclusion Framework for Visually Interpretable and Uncertainty-Guided Out-of-Distribution Detection
- Training-Free Diffusion Priors for Text-to-Image Generation via Optimization-based Visual Inversion
- RankOOD -- Class Ranking-based Out-of-Distribution Detection
- Latent-space metrics for Complex-Valued VAE out-of-distribution detection under radar clutter
- SupLID: Geometrical Guidance for Out-of-Distribution Detection in Semantic Segmentation
- CroTad: A Contrastive Reinforcement Learning Framework for Online Trajectory Anomaly Detection
- Self-Supervised Adversarial Example Detection by Disentangled Representation
- Exploiting Inter-Sample Information for Long-tailed Out-of-Distribution Detection
- SNAP: Low-Latency Test-Time Adaptation with Sparse Updates
- TSRE: Channel-Aware Typical Set Refinement for Out-of-Distribution Detection
- Inverse Electromagnetic Scattering for Doubly-Connected Cylinders using Convolutional Neural Networks
- BootOOD: Self-Supervised Out-of-Distribution Detection via Synthetic Sample Exposure under Neural Collapse
- Multi-Loss Sub-Ensembles for Accurate Classification with Uncertainty Estimation
- A Simple Fix to Mahalanobis Distance for Improving Near-OOD Detection
- Risk Management Framework for Machine Learning Security
- Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
- A Systematic Analysis of Out-of-Distribution Detection Under Representation and Training Paradigm Shifts
- DeepDefense: Layer-Wise Gradient-Feature Alignment for Building Robust Neural Networks
- Taming Object Hallucinations with Verified Atomic Confidence Estimation
- Brittle Features May Help Anomaly Detection
- CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features
- Automatic Grid Updates for Kolmogorov-Arnold Networks using Layer Histograms
- ClusterMine: Robust Label-Free Visual Out-Of-Distribution Detection via Concept Mining from Text Corpora
- Relative Energy Learning for LiDAR Out-of-Distribution Detection
- Deep learning models are vulnerable, but adversarial examples are even more vulnerable
- Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
- Measuring Aleatoric and Epistemic Uncertainty in LLMs: Empirical Evaluation on ID and OOD QA Tasks
- GAFD-CC: Global-Aware Feature Decoupling with Confidence Calibration for OOD Detection
- Uncertainty Guided Online Ensemble for Non-stationary Data Streams in Fusion Science
- Perturbations in the Orthogonal Complement Subspace for Efficient Out-of-Distribution Detection
- Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
- Covariance Last-Layer Ensembles: Function-Space Diversity for Efficient Uncertainty Quantification
- On the detection of Out-Of-Distribution samples in Multiple Instance Learning
- Identifying Untrustworthy Predictions in Neural Networks by Geometric\n Gradient Analysis
- Probabilistic Trust Intervals for Out of Distribution Detection
- Level, Sharpness, and Corpus: Why Zero-Shot OOD Detector Rankings Do Not Transfer
- Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- Novelty Detection via Robust Variational Autoencoding
- Lifelong Machine Learning with Deep Streaming Linear Discriminant Analysis
- Distilling Causal Effect of Data in Class-Incremental Learning
- GNNGuard: Defending Graph Neural Networks against Adversarial Attacks
- ToxScreen: Detecting Whether an LLM Has Been Poisoned
- Anomalous Example Detection in Deep Learning: A Survey
- On Last-Layer Algorithms for Classification: Decoupling Representation from Uncertainty Estimation
- Fine-grained Uncertainty Modeling in Neural Networks
- Generalized Out-of-Distribution Detection: A Survey
- Contrastive Training for Improved Out-of-Distribution Detection
- Trust Issues: Uncertainty Estimation Does Not Enable Reliable OOD Detection On Medical Tabular Data
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted Instances
- When Rule Violations Are Rare: Chimera Training for Logical Anomaly Detection
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen Attacks
- Robust Anomaly Detection for Particle Physics Using Multi-Background Representation Learning
- One Versus all for deep Neural Network Incertitude (OVNNI) quantification
- SegmentMeIfYouCan: A Benchmark for Anomaly Segmentation
- Detecting Out-of-Distribution Examples with In-distribution Examples and Gram Matrices
- Energy-based Out-of-distribution Detection
- Likelihood Landscapes: A Unifying Principle Behind Many Adversarial Defenses
- Likelihood Regret: An Out-of-Distribution Detection Score For Variational Auto-encoder
- MOOD: Multi-level Out-of-distribution Detection
- What Are Bayesian Neural Network Posteriors Really Like?
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Out-of-Distribution Detection for Safety Assurance of AI and Autonomous Systems
- FINE Samples for Learning with Noisy Labels
- Revisiting Logit Distributions for Reliable Out-of-Distribution Detection
- Towards Strong Certified Defense with Universal Asymmetric Randomization
- Beyond Binary Out-of-Distribution Detection: Characterizing Distributional Shifts with Multi-Statistic Diffusion Trajectories
- Learning After Model Deployment
- GOOD: Training-Free Guided Diffusion Sampling for Out-of-Distribution Detection
- Benchmarking Out-of-Distribution Detection for Plankton Recognition: A Systematic Evaluation of Advanced Methods in Marine Ecological Monitoring
- Deep Learning-Based Autonomous Driving Systems: A Survey of Attacks and Defenses
- Dissecting Mahalanobis: How Feature Geometry and Normalization Shape OOD Detection
- Further Analysis of Outlier Detection with Deep Generative Models
- PriorGuide: Test-Time Prior Adaptation for Simulation-Based Inference
- Shifting Transformation Learning for Out-of-Distribution Detection
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- A Multi-dimensional Semantic Surprise Framework Based on Low-Entropy Semantic Manifolds for Fine-Grained Out-of-Distribution Detection
- KoALA: KL-L0 Adversarial Detector via Label Agreement
- Local Background Features Matter in Out-of-Distribution Detection
- Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions
- A Machine Learning Perspective on Automated Driving Corner Cases
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
- Out-of-domain Detection for Natural Language Understanding in Dialog Systems
- Robust Canonicalization through Bootstrapped Data Re-Alignment
- EigenScore: OOD Detection using Covariance in Diffusion Models
- Out-of-Distribution Detection in LiDAR Semantic Segmentation Using Epistemic Uncertainty from Hierarchical GMMs
- Medix: Out-of-Distribution Detection from Unlabeled Wild Data via Robust Gradient Statistics
- Valid Stopping for LLM Generation via Empirical Dynamic Formal Lift
- Mysteries of the Deep: Role of Intermediate Representations in Out of Distribution Detection
- Human Texts Are Outliers: Detecting LLM-generated Texts via Out-of-distribution Detection
- Out-of-Distribution Detection from Small Training Sets using Bayesian Neural Network Classifiers
- Attack of the Tails: Yes, You Really Can Backdoor Federated Learning
- SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics
- A Statistical Method for Attack-Agnostic Adversarial Attack Detection with Compressive Sensing Comparison
- Can multi-label classification networks know what they don't know?
- Can Molecular Foundation Models Know What They Don't Know? A Simple Remedy with Preference Optimization
- Meta Learning Low Rank Covariance Factors for Energy-Based Deterministic\n Uncertainty
- Calibration of Model Uncertainty for Dropout Variational Inference
- Multidimensional Uncertainty Quantification via Optimal Transport
- HyperCore: Coreset Selection under Noise via Hypersphere Models
- Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
- Exploring the Limits of Out-of-Distribution Detection
- ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research
- Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances
- Analysis of Confident-Classifiers for Out-of-distribution Detection
- PGTuner: An Efficient Framework for Automatic and Transferable Configuration Tuning of Proximity Graphs
- SSD: A Unified Framework for Self-Supervised Outlier Detection
- Learning to Separate Clusters of Adversarial Representations for Robust Adversarial Detection
- Probabilistic Runtime Verification, Evaluation and Risk Assessment of Visual Deep Learning Systems
- Text Meets Topology: Rethinking Out-of-distribution Detection in Text-Rich Networks
- Robustness Feature Adapter for Efficient Adversarial Training
- Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
- AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
- Long-Tailed Out-of-Distribution Detection with Refined Separate Class Learning
- Deep Learning and Traffic Classification: Lessons learned from a commercial-grade dataset with hundreds of encrypted and zero-day applications
- Generalizing Neural Networks by Reflecting Deviating Data in Production
- DOI: Divergence-based Out-of-Distribution Indicators via Deep Generative Models
- Dynamic Aware: Adaptive Multi-Mode Out-of-Distribution Detection for Trajectory Prediction in Autonomous Vehicles
- An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
- Out of Distribution Detection in Self-adaptive Robots with AI-powered Digital Twins
- Dataset Inference: Ownership Resolution in Machine Learning
- Sharpness-Aware Geometric Defense for Robust Out-Of-Distribution Detection
- MOS: Towards Scaling Out-of-distribution Detection for Large Semantic Space
- Logit Mixture Outlier Exposure for Fine-grained Out-of-Distribution Detection
- Generalized ODIN: Detecting Out-of-distribution Image without Learning from Out-of-distribution Data
- Self-Supervised Training Enhances Online Continual Learning
- A Critical Evaluation of Open-World Machine Learning
- Prompt Optimization Meets Subspace Representation Learning for Few-shot Out-of-Distribution Detection
- Tackling the Noisy Elephant in the Room: Label Noise-robust Out-of-Distribution Detection via Loss Correction and Low-rank Decomposition
- DCV-ROOD Evaluation Framework: Dual Cross-Validation for Robust Out-of-Distribution Detection
- Prior Distribution and Model Confidence
- Polysemantic Dropout: Conformal OOD Detection for Specialized LLMs
- ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
- Exploring Vicinal Risk Minimization for Lightweight Out-of-Distribution Detection
- LiBRe: A Practical Bayesian Approach to Adversarial Detection
- Mean-Field Approximation to Gaussian-Softmax Integral with Application to Uncertainty Estimation
- Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare
- Practical Evaluation of Out-of-Distribution Detection Methods for Image\n Classification
- Entropy-Based Non-Invasive Reliability Monitoring of Convolutional Neural Networks
- Activation Subspaces for Out-of-Distribution Detection
- Multi-Method Ensemble for Out-of-Distribution Detection
- Bigeminal Priors Variational auto-encoder
- E-Stitchup: Data Augmentation for Pre-Trained Embeddings
- Understanding Classifier Mistakes with Generative Models
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
- Probability Density from Latent Diffusion Models for Out-of-Distribution Detection
- Likelihood Ratios for Out-of-Distribution Detection
- ReLATE+: Unified Framework for Adversarial Attack Detection, Classification, and Resilient Model Selection in Time-Series Classification
- GraN: An Efficient Gradient-Norm Based Detector for Adversarial and\n Misclassified Examples
- Principled Detection of Hallucinations in Large Language Models via Multiple Testing
- Uncertainty Estimation Using a Single Deep Deterministic Neural Network
- Beyond Turing: Memory-Amortized Inference as a Foundation for Cognitive Computation
- Empirical Evidences for the Effects of Feature Diversity in Open Set Recognition and Continual Learning
- SNAP-UQ: Self-supervised Next-Activation Prediction for Single-Pass Uncertainty in TinyML
- Autoregressive Models: What Are They Good For?
- M3OOD: Automatic Selection of Multimodal OOD Detectors
- On the Sample Complexity of Adversarial Multi-Source PAC Learning
- Retrieval-Augmented Prompt for OOD Detection
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- Out-of-Distribution Detection using Counterfactual Distance
- OpenHAIV: A Framework Towards Practical Open-World Learning
- Prediction of Survival Outcomes under Clinical Presence Shift: A Joint Neural Network Architecture
- Generalized Few-Shot Out-of-Distribution Detection
- Increasing Trustworthiness of Deep Neural Networks via Accuracy Monitoring
- Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
- Improved Robustness to Open Set Inputs via Tempered Mixup
- Pseudo-label Induced Subspace Representation Learning for Robust Out-of-Distribution Detection
- A Unified Plug-and-Play Framework for Effective Data Denoising and Robust Abstention
- Simple Methods Defend RAG Systems Well Against Real-World Attacks
- TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
- Accumulative Poisoning Attacks on Real-time Data
- A Simple and Effective Method for Uncertainty Quantification and OOD Detection
- Uncertainty-Aware Likelihood Ratio Estimation for Pixel-Wise Out-of-Distribution Detection
- BOOD: Boundary-based Out-Of-Distribution Data Generation
Related