Foundations of Interpretable Models
2025/08/01 by Barbiero, Pietro, Zarlenga, Mateo Espinosa, Termine, Alberto +2 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE)
paper · doi:10.48550/arxiv.2508.00545
Abstract
We argue that existing definitions of interpretability are not actionable in that they fail to inform users about general, sound, and robust interpretable model design. This makes current interpretability research fundamentally ill-posed. To address this issue, we propose a definition of interpretability that is general, simple, and subsumes existing informal notions within the interpretable AI community. We show that our definition is actionable, as it directly reveals the foundational properties, underlying assumptions, principles, data structures, and architectural features necessary for designing interpretable models. Building on this, we propose a general blueprint for designing interpretable models and introduce the first open-sourced library with native support for interpretable data structures and processes.
Citations
- Interpretable Hierarchical Concept Reasoning through Attention-Guided Graph Learning
- We Can't Understand AI Using our Existing Vocabulary
- A Complexity-Based Theory of Compositionality
- Causal Concept Graph Models: Beyond Causal Opacity in Deep Learning
- Beyond Concept Bottleneck Models: How to Make Black Boxes Intervenable?
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Representation Engineering: A Top-Down Approach to AI Transparency
- Learning to Receive Help: Intervention-Aware Concept Embedding Models
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Interpretability is in the Mind of the Beholder: A Causal Framework for Human-interpretable Representation Learning
- Selective Concept Models: Permitting Stakeholder Customisation at Test-Time
- The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
- Human Uncertainty in Concept-Based AI Systems
- Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
- Discovering Latent Knowledge in Language Models Without Supervision
- Training language models to follow instructions with human feedback
- Promises and Pitfalls of Black-Box Concept Learning Models
- Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
- Concept Bottleneck Models
- Shortcut learning in deep neural networks
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI
- On Completeness-aware Concept-Based Explanations in Deep Neural Networks
- Explaining Classifiers with Causal Concept Effect (CaCE)
- Towards Robust Interpretability with Self-Explaining Neural Networks
- Seven Sketches in Compositionality: An Invitation to Applied Category\n Theory
- Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
- Generalization in Deep Learning
- Exploring Generalization in Deep Learning
- Explanation in Artificial Intelligence: Insights from the Social\n Sciences
- Explanation in artificial intelligence: Insights from the social sciences
- A Unified Approach to Interpreting Model Predictions
- Axiomatic Attribution for Deep Networks
- Distilling the Knowledge in a Neural Network
- Representation Learning: A Review and New Perspectives
- Representation Learning: A Review and New Perspectives
- I.—COMPUTING MACHINERY AND INTELLIGENCE
- The magical number seven, plus or minus two: Some limits on our capacity for processing information.
Cited by
Related