Towards A Rigorous Science of Interpretable Machine Learning
2017/02/28 by Doshi-Velez, Finale, Kim, Been · 137 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
paper · doi:10.48550/arxiv.1702.08608
Abstract
As machine learning systems become ubiquitous, there has been a surge of interest in interpretable machine learning: systems that provide explanation for their outputs. These explanations are often used to qualitatively assess other criteria such as safety or non-discrimination. However, despite the interest in interpretability, there is very little consensus on what interpretable machine learning is and how it should be measured. In this position paper, we first define interpretability and describe when interpretability is needed (and when it is not). Next, we suggest a taxonomy for rigorous evaluation and expose open questions towards a more rigorous science of interpretable machine learning.
Cited by
- Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
- When Algorithms Manage Humans: A Double Machine Learning Approach to Estimating Nonlinear Effects of Algorithmic Control on Gig Worker Performance and Wellbeing
- Toward Human-Centered Multi-Agent Systems: Integrating Cognition, Culture, Values, and Cooperation in AI Agents
- Beyond Adoption Intention How Trust in Augmented Analytics Relates to Perceived Decision Quality Among Non-Technical BI Users
- JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI
- Explainability Methods for Hardware Trojan Detection: A Systematic Comparison
- Q-A3C2: Quantum Reinforcement Learning with Time-Series Dynamic Clustering for Adaptive ETF Stock Selection
- From artificial to organic: Rethinking the roots of intelligence for digital health
- Block-Recurrent Dynamics in Vision Transformers
- Cluster-Based Generalized Additive Models Informed by Random Fourier Features
- Explanation Beyond Intuition: A Testable Criterion for Inherent Explainability
- The Social Responsibility Stack: A Control-Theoretic Architecture for Governing Socio-Technical AI
- The Agony of Opacity: Foundations for Reflective Interpretability in AI-Mediated Mental Health Support
- Where is the Watermark? Interpretable Watermark Detection at the Block Level
- A Neuro-Symbolic Framework for Accountability in Public-Sector AI
- Human-AI Collaboration Mechanism Study on AIGC Assisted Image Production for Special Coverage
- Explainable Artificial Intelligence for Economic Time Series: A Comprehensive Review and a Systematic Taxonomy of Methods and Concepts
- Interaction Tensor SHAP
- Back to the Baseline: Examining Baseline Effects on Explainability Metrics
- Reinforcement Learning in Financial Decision Making: A Systematic Review of Performance, Challenges, and Implementation Strategies
- A Geometric Unification of Concept Learning with Concept Cones
- Interpretive Efficiency: Information-Geometric Foundations of Data Usefulness
- Modular Jets for Supervised Pipelines: Diagnosing Mirage vs Identifiability
- MASE: Interpretable NLP Models via Model-Agnostic Saliency Estimation
- Menta: A Small Language Model for On-Device Mental Health Prediction
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model Predictions
- Faster Verified Explanations for Neural Networks
- Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
- The Directed Prediction Change - Efficient and Trustworthy Fidelity Assessment for Local Feature Attribution Methods
- Lean 5.0: A Predictive, Human-AI, and Ethically Grounded Paradigm for Construction Management
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- SG-OIF: A Stability-Guided Online Influence Framework for Reliable Vision Data
- Accuracy is Not Enough: Poisoning Interpretability in Federated Learning via Color Skew
- Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for Explanations
- Extremal Contours: Gradient-driven contours for compact visual attribution
- MACIE: Multi-Agent Causal Intelligence Explainer for Collective Behavior Understanding
- Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring
- Flexible Concept Bottleneck Model
- Function Based Isolation Forest (FuBIF): A Unifying Framework for Interpretable Isolation-Based Anomaly Detection
- Search Is Not Retrieval: Decoupling Semantic Matching from Contextual Assembly in RAG
- Efficacy Analysis in Clinical Trials: A Comprehensive Review of Statistical and Machine Learning Approaches
- T-FIX: Text-Based Explanations with Features Interpretable to eXperts
- Human Resource Management and AI: A Contextual Transparency Database
- Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
- Imitation Learning in the Deep Learning Era: A Novel Taxonomy and Recent Advances
- Fair and Explainable Credit-Scoring under Concept Drift: Adaptive Explanation Frameworks for Evolving Populations
- Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach
- The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold
- On the Emergence of Induction Heads for In-Context Learning
- Before the Clinic: Transparent and Operable Design Principles for Healthcare AI
- Interpreting LLMs as Credit Risk Classifiers: Do Their Feature Explanations Align with Classical ML?
- Teaching Machine Learning to Software Engineers
- What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
- Bridging Accuracy and Interpretability: Deep Learning with XAI for Breast Cancer Detection
- GroupSHAP-Guided Integration of Financial News Keywords and Technical Indicators for Stock Price Prediction
- Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges
- K-DAREK: Distance Aware Error for Kurkova Kolmogorov Networks
- Soppia: A Structured Prompting Framework for the Proportional Assessment of Non-Pecuniary Damages in Personal Injury Cases
- Framework for Machine Evaluation of Reasoning Completeness in Large Language Models For Classification Tasks
- Interpretable machine learning for identifying individual-specific cardiogram signatures
- When LRP Diverges from Leave-One-Out in Transformers
- Empowering Decision Trees via Shape Function Branching
- The Impact of Concept Explanations and Interventions on Human-Machine Collaboration
- Programmatic Representation Learning with Language Models
- DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
- On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- Making Power Explicable in AI: Analyzing, Understanding, and Redirecting Power to Operationalize Ethics in AI Technical Practice
- From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening
- Training Feature Attribution for Vision Models
- ClauseLens: Clause-Grounded, CVaR-Constrained Reinforcement Learning for Trustworthy Reinsurance Pricing
- Deep Neural Networks Inspired by Differential Equations
- Post-hoc Stochastic Concept Bottleneck Models
- Faithful and Interpretable Explanations for Complex Ensemble Time Series Forecasts using Surrogate Models and Forecastability Analysis
- Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
- Cluster Paths: Navigating Interpretability in Neural Networks
- QGraphLIME - Explaining Quantum Graph Neural Networks
- Trust in Transparency: How Explainable AI Shapes User Perceptions
- Tail-Safe Hedging: Explainable Risk-Sensitive Reinforcement Learning with a White-Box CBF--QP Safety Layer in Arbitrage-Free Markets
- Reconsidering Requirements Engineering: Human-AI Collaboration in AI-Native Software Development
- Combining Large Language Models and Gradient-Free Optimization for Automatic Control Policy Synthesis
- MetaSynth: Multi-Agent Metadata Generation from Implicit Feedback in Black-Box Systems
- ACT: Agentic Classification Tree
- ACE: Adapting sampling for Counterfactual Explanations
- PET: Preference Evolution Tracking with LLM-Generated Explainable Distribution
- Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
- MindCraft: How Concept Trees Take Shape In Deep Models
- CausalKANs: interpretable treatment effect estimation with Kolmogorov-Arnold networks
- VizGen: Data Exploration and Visualization from Natural Language via a Multi-Agent AI Architecture
- Outlier Detection in Plantar Pressure: Human-Centered Comparison of Statistical Parametric Mapping and Explainable Machine Learning
- X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
- Robust, Observable, and Evolvable Agentic Systems Engineering: A Principled Framework Validated via the Fairy GUI Agent
- Achieving Fair Skin Lesion Detection through Skin Tone Normalization and Channel Pruning
- Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
- Desde la caja negra a la comprensión: fortaleciendo el vínculo modelo-fenómeno mediante inteligencia artificial explicable
- From moral panic to pragmatic governance: reframing AI’s societal impacts in employment, education, and ethics
- Explainable Graph Neural Networks: Understanding Brain Connectivity and Biomarkers in Dementia
- Truth Without Comprehension: A BlueSky Agenda for Steering the Fourth Mathematical Crisis
- Philosophy-informed Machine Learning
- A Scenario-Driven Cognitive Approach to Next-Generation AI Memory
- Clarifying Model Transparency: Interpretability versus Explainability in Deep Learning with MNIST and IMDB Examples
- NuGraph2 with Explainability: Post-hoc Explanations for Geometric Neural Network Predictions
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"
- Interpretability as Alignment: Making Internal Understanding a Design Principle
- An Interpretable Deep Learning Model for General Insurance Pricing
- Temporal Counterfactual Explanations of Behaviour Tree Decisions
- REMI: A Novel Causal Schema Memory Architecture for Personalized Lifestyle Recommendation Agents
- From Eigenmodes to Proofs: Integrating Graph Spectral Operators with Symbolic Interpretable Reasoning
- An Approach to Grounding AI Model Evaluations in Human-derived Criteria
- Sparse Autoencoder Neural Operators: Model Recovery in Function Spaces
- Explainability-Driven Dimensionality Reduction for Hyperspectral Imaging
- Theory Foundation of Physics-Enhanced Residual Learning
- Do Cognitively Interpretable Reasoning Traces Improve LLM Performance?
- Cross-Attention Multimodal Fusion for Breast Cancer Diagnosis: Integrating Mammography and Clinical Data with Explainability
- Interestingness First Classifiers
- A Dynamical Systems Framework for Reinforcement Learning Safety and Robustness Verification
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- The AI-Fraud Diamond: A Novel Lens for Auditing Algorithmic Deception
- Documenting Deployment with Fabric: A Repository of Real-World AI Governance
- LumiMAS: A Comprehensive Framework for Real-Time Monitoring and Enhanced Observability in Multi-Agent Systems
- A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond
- RealAC: A Domain-Agnostic Framework for Realistic and Actionable Counterfactual Explanations
- A Unified Evaluation Framework for Multi-Annotator Tendency Learning
- Your Recourse, My Loss? Algorithmic Recourse under Shared Constraints
- Beyond Technocratic XAI: The Who, What & How in Explanation Design
- A Moral Agency Framework for Legitimate Integration of AI in Bureaucracies
- Attribution Explanations for Deep Neural Networks: A Theoretical Perspective
- MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
- Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems
- A Feature Engineering Approach for Business Impact-Oriented Failure Detection in Distributed Instant Payment Systems
- I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
- Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts
- Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust
- Toward using explainable data-driven surrogate models for treating performance-based seismic design as an inverse engineering problem
- CTG-Insight: A Multi-Agent Interpretable LLM Framework for Cardiotocography Analysis and Classification
Related