An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations
2025/05/22 by S. J. Park, Park, Seonghwan, Jueun Mun +5
Computer Science · #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.2505.16705
openalex publication_date 2025/05/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Abstract
Concept bottleneck models (CBMs) ensure interpretability by decomposing predictions into human interpretable concepts. Yet the annotations used for training CBMs that enable this transparency are often noisy, and the impact of such corruption is not well understood. In this study, we present the first systematic study of noise in CBMs and show that even moderate corruption simultaneously impairs prediction performance, interpretability, and the intervention effectiveness. Our analysis identifies a susceptible subset of concepts whose accuracy declines far more than the average gap between noisy and clean supervision and whose corruption accounts for most performance loss. To mitigate this vulnerability we propose a two-stage framework. During training, sharpness-aware minimization stabilizes the learning of noise-sensitive concepts. During inference, where clean labels are unavailable, we rank concepts by predictive entropy and correct only the most uncertain ones, using uncertainty as a proxy for susceptibility. Theoretical analysis and extensive ablations elucidate why sharpness-aware training confers robustness and why uncertainty reliably identifies susceptible concepts, providing a principled basis that preserves both interpretability and resilience in the presence of noise.
Citations
- Concept Bottleneck Large Language Models
- VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
- Semi-supervised Concept Bottleneck Models
- Stochastic Concept Bottleneck Models
- Why is SAM Robust to Label Noise?
- Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Probabilistic Interpretations
- Auxiliary Losses for Learning Generalizable Concept-based Models
- Interpreting Pretrained Language Models via Concept Bottlenecks
- Coarse-to-Fine Concept Bottleneck Models
- A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations
- A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations
- Label-Free Concept Bottleneck Models
- A Closer Look at the Intervention Procedure of Concept Bottleneck Models
- Understanding and Enhancing Robustness of Concept-based Models
- Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification
- Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
- Concept Bottleneck Model with Additional Unsupervised Concepts
- Towards Understanding Deep Learning from Noisy Labels with Small-Loss Criterion
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges
- Tackling Instance-Dependent Label Noise via a Universal Probabilistic Model
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- Debiasing Concept-based Explanations with Causal Analysis
- Concept Bottleneck Models
- Identifying Mislabeled Data using the Area Under the Margin Ranking
- Symmetric Cross Entropy for Robust Learning with Noisy Labels
- Combating Label Noise in Deep Learning Using Abstention
- Unsupervised Label Noise Modeling and Loss Correction
- Dimensionality-Driven Learning with Noisy Labels
- Co-teaching: Robust Training of Deep Neural Networks with Extremely Noisy Labels
- mixup: Beyond Empirical Risk Minimization
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- Zero-Shot Learning -- A Comprehensive Evaluation of the Good, the Bad and the Ugly
- Attention Is All You Need
- Regularizing Neural Networks by Penalizing Confident Output Distributions
- A Tutorial on Kernel Density Estimation and Recent Advances
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
- Pruning Filters for Efficient ConvNets
- Deep Residual Learning for Image Recognition
- Rethinking the Inception Architecture for Computer Vision
- A sequential algorithm for training text classifiers
Related