The Mythos of Model Interpretability
2018/06/01 by Zachary C. Lipton · 140 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning #Machine Learning and Data Classification
paper · doi:10.1145/3233231
Abstract
In machine learning, the concept of interpretability is both important and slippery.
Citations
Cited by
- Explanation in artificial intelligence: Insights from the social sciences
- Local Additive Feature Attribution: A Mathematical Taxonomy and Reporting Checklist
- Automatic Construction of Clinical Scoring Systems with LLM Agents
- Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm
- Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft
- Machine Learning Approaches for Improved Scalability of Metallic Magnetic Calorimeters
- A modular state-space model of human perception, cognition, and decision dynamics
- IMEX Interaction-Based Model Explanation
- Spectral biclustering-driven scalability for post-hoc explainability in recommender systems
- People Can Accurately Predict Behavior of Complex Algorithms That Are Available, Compact, and Aligned
- Cosine capital: Large language models and the embedding of all things
- Explainable AI in Healthcare: to Explain, to Predict, or to Describe?
- Position: We Need An Algorithmic Understanding of Generative AI
- Legally-Informed Explainable AI
- Better sampling in explanation methods can prevent dieselgate-like deception
- Learn to Explain Efficiently via Neural Logic Inductive Learning
- ExAD: An Ensemble Approach for Explanation-based Adversarial Detection
- Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
- EvoXplain: When Machine Learning Models Agree on Predictions but Disagree on Why -- Measuring Mechanistic Multiplicity Across Training Runs
- Generative AI for Requirements Engineering: A Systematic Literature Review
- Modeling Missing Data in Clinical Time Series with RNNs
- Why an Android App is Classified as Malware? Towards Malware Classification Interpretation
- Coherent without Grounding, Grounded without Success: The Bidirectional Coherence Paradox in Artificial Epistemic Agents
- ML-LOO: Detecting Adversarial Examples with Feature Attribution
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Learning Less-Overlapping Representations
- Exploring Interpretable LSTM Neural Networks over Multi-Variable Data
- On the Theoretical Foundation of Sparse Dictionary Learning in Mechanistic Interpretability
- MASE: Interpretable NLP Models via Model-Agnostic Saliency Estimation
- MOTIF-RF: Multi-template On-chip Transformer Synthesis Incorporating Frequency-domain Self-transfer Learning for RFIC Design Automation
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model Predictions
- DRLViz: Understanding Decisions and Memory in Deep Reinforcement\n Learning
- Softly Symbolifying Kolmogorov-Arnold Networks
- What Clinicians Want: Contextualizing Explainable Machine Learning for Clinical End Use
- Semantic Explanations of Predictions
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Bridging Philosophy and Machine Learning: A Structuralist Framework for Classifying Neural Network Representations
- Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for Explanations
- Constructing and Evaluating an Explainable Model for COVID-19 Diagnosis from Chest X-rays
- Can You Explain That, Better? Comprehensible Text Analytics for SE Applications
- Efficient Search for Diverse Coherent Explanations
- Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring
- A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning
- Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
- Trust Considerations for Explainable Robots: A Human Factors Perspective
- QiNN-QJ: A Quantum-inspired Neural Network with Quantum Jump for Multimodal Sentiment Analysis
- "Show Me You Comply... Without Showing Me Anything": Zero-Knowledge Software Auditing for AI-Enabled Systems
- Interpretable and Pedagogical Examples
- Regularizing Explanations in Bayesian Convolutional Neural Networks
- Methods for Interpreting and Understanding Deep Neural Networks
- DANCE: Enhancing saliency maps using decoys
- Applications of Psychological Science for Actionable Analytics
- Interpretation of Neural Networks is Fragile
- Utilising Deep Learning and Genome Wide Association Studies for\n Epistatic-Driven Preterm Birth Classification in African-American Women
- State of the Art in Fair ML: From Moral Philosophy and Legislation to\n Fair Classifiers
- ConvNets and ImageNet Beyond Accuracy: Understanding Mistakes and\n Uncovering Biases
- Machine learning approaches for interpretable antibody property prediction using structural data
- Clusters in Explanation Space: Inferring disease subtypes from model explanations
- Deep Neural Networks for Choice Analysis: Architectural Design with Alternative-Specific Utility Functions
- Shapley Interpretation and Activation in Neural Networks
- Post-hoc Stochastic Concept Bottleneck Models
- Cluster Paths: Navigating Interpretability in Neural Networks
- Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
- Trust in Transparency: How Explainable AI Shapes User Perceptions
- A Human-Grounded Evaluation of SHAP for Alert Processing
- Causal Interpretability for Machine Learning -- Problems, Methods and Evaluation
- Contrastive Explanations with Local Foil Trees
- MindCraft: How Concept Trees Take Shape In Deep Models
- When Can AI Models Explain Learning? Validity Criteria for AI as Cognitive Models in Education
- (Sometimes) Less is More: Mitigating the Complexity of Rule-based Representation for Interpretable Classification
- LAVA: Explainability for Unsupervised Latent Embeddings
- Explainable time series tweaking via irreversible and reversible temporal transformations
- INCLAIR: Inception-Based Longitudinal Clinical Anomaly Detection with Informed Reasoning
- Archival Paradata and Artificial Intelligence in Archaeology
- Train, Diagnose and Fix: Interpretable Approach for Fine-grained Action Recognition
- Relating Input Concepts to Convolutional Neural Network Decisions
- Challenging common interpretability assumptions in feature attribution explanations
- Learning Interpretable Concept-Based Models with Human Feedback
- On the Philosophical Naivety of Engineers in the Age of Machine Learning
- Independently Interpretable Lasso: A New Regularizer for Sparse Regression with Uncorrelated Variables
- Robust and Stable Black Box Explanations
- Efficient & Correct Predictive Equivalence for Decision Trees
- Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning
- TensorFuzz: Debugging Neural Networks with Coverage-Guided Fuzzing
- Towards a Transparent and Interpretable AI Model for Medical Image Classifications
- Explaining Question Answering Models through Text Generation
- Transparent and Fair Profiling in Employment Services: Evidence from Switzerland
- Automated Rationale Generation: A Technique for Explainable AI and its Effects on Human Perceptions
- It’s not a bug, it’s a feature: How AI experts and data scientists account for the opacity of algorithms
- Clarifying Model Transparency: Interpretability versus Explainability in Deep Learning with MNIST and IMDB Examples
- Progressive Disclosure: Designing for Effective Transparency
- NuGraph2 with Explainability: Post-hoc Explanations for Geometric Neural Network Predictions
- NBDT: Neural-Backed Decision Trees
- Explainability of CNN Based Classification Models for Acoustic Signal
- Interpretability as Alignment: Making Internal Understanding a Design Principle
- Transparency of medical artificial intelligence systems
- Breaking SafetyCore: Exploring the Risks of On-Device AI Deployment
- Direct Network Transfer: Transfer Learning of Sentence Embeddings for Semantic Similarity
- From Eigenmodes to Proofs: Integrating Graph Spectral Operators with Symbolic Interpretable Reasoning
- DRLViz: Understanding Decisions and Memory in Deep Reinforcement Learning
- Explainable Artificial Intelligence Approaches: A Survey
- TIME: A Transparent, Interpretable, Model-Adaptive and Explainable Neural Network for Dynamic Physical Processes
- A case study of forensic psychiatry experts' reports analysis through large language models
- Improving the Interpretability of Deep Neural Networks with Knowledge Distillation
- Extending LIME for Business Process Automation
- Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning
- Individualized and Interpretable Sleep Forecasting via a Two-Stage Adaptive Spatial-Temporal Model
- Detecting Statistical Interactions from Neural Network Weights
- Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding
- Scalable Rule-Based Representation Learning for Interpretable Classification
- Interestingness First Classifiers
- The AI-Fraud Diamond: A Novel Lens for Auditing Algorithmic Deception
- Goal-Directedness is in the Eye of the Beholder
- How can we trust opaque systems? Criteria for robust explanations in XAI
- Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks
- Explainable artificial intelligence (XAI), the goodness criteria and the\n grasp-ability test
- The Wrath of KAN: Enabling Fast, Accurate, and Transparent Emulation of the Global 21 cm Cosmology Signal
- Artificial Intelligence and Black‐Box Medical Decisions: <i>Accuracy versus Explainability</i>
- To Explain Or Not To Explain: An Empirical Investigation Of AI-Based Recommendations On Social Media Platforms
- A Moral Agency Framework for Legitimate Integration of AI in Bureaucracies
- The Robust Manifold Defense: Adversarial Training using Generative Models
- AI with Symbolic Empathy: Shannon-Neumann Insight Guided Logic
- Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems
- Discovering Conditionally Salient Features with Statistical Guarantees
- Toward using explainable data-driven surrogate models for treating performance-based seismic design as an inverse engineering problem
- Machine Learning Pipeline for Software Engineering: A Systematic Literature Review
- Your Model Is Unfair, Are You Even Aware? Inverse Relationship Between Comprehension and Trust in Explainability Visualizations of Biased ML Models
- Discrete Choice Analysis with Machine Learning Capabilities
- Teaching the Teacher: Improving Neural Network Distillability for Symbolic Regression via Jacobian Regularization
- Increasing the Interpretability of Recurrent Neural Networks Using Hidden Markov Models
- Explaining hyperspectral imaging based plant disease identification: 3D CNN and saliency maps
- Cognitive Psychology for Deep Neural Networks: A Shape Bias Case Study
- Tree Space Prototypes: Another Look at Making Tree Ensembles Interpretable
- LEAFAGE: Example-based and Feature importance-based Explanationsfor Black-box ML models
- SIFOTL: A Principled, Statistically-Informed Fidelity-Optimization Method for Tabular Learning
- What Does Explainable AI Really Mean? A New Conceptualization of Perspectives
- How can deep learning advance computational modeling of sensory information processing?
- Interpreting search result rankings through intent modeling
- Mining Electronic Health Records: A Survey
- Explaining Groups of Points in Low-Dimensional Representations
- The Performance of Different Artificial Intelligence Models in Predicting Breast Cancer among Individuals Having Type 2 Diabetes Mellitus. [europepmc]
- Improving reference prioritisation with PICO recognition. [europepmc]
- Deep learning for image-based large-flowered chrysanthemum cultivar recognition. [europepmc]
- The role of artificial intelligence in achieving the Sustainable Development Goals. [europepmc]
- Interpretable clinical prediction via attention-based neural network. [europepmc]
- Predicting Early Warning Signs of Psychotic Relapse From Passive Sensing Data: An Approach Using Encoder-Decoder Neural Networks. [europepmc]
- Validation of a Machine Learning Model to Predict Childhood Lead Poisoning. [europepmc]
- Illuminating the Black Box: Interpreting Deep Neural Network Models for Psychiatric Research. [europepmc]
- Deep Learning in LncRNAome: Contribution, Challenges, and Perspectives. [europepmc]
- An Interpretation Architecture for Deep Learning Models with the Application of COVID-19 Diagnosis. [europepmc]
- Towards the Interpretability of Machine Learning Predictions for Medical Applications Targeting Personalised Therapies: A Cancer Case Survey. [europepmc]
- Population Preferences for Performance and Explainability of Artificial Intelligence in Health Care: Choice-Based Conjoint Survey. [europepmc]
- Using Machine Learning to Predict Complications in Pregnancy: A Systematic Review. [europepmc]
- Rapid SERS identification of methicillin-susceptible and methicillin-resistant Staphylococcus aureus via aptamer recognition and deep learning. [europepmc]
- Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. [europepmc]
- Trusting our machines: validating machine learning models for single-molecule transport experiments. [europepmc]
- Opening the black box: interpretable machine learning for predictor finding of metabolic syndrome. [europepmc]
- Designing the ultrasonic treatment of nanoparticle-dispersions via machine learning. [europepmc]
- Interpretable machine learning methods for predictions in systems biology from omics data. [europepmc]
- Explainable AI: A review of applications to neuroimaging data. [europepmc]
- Interpretability of Clinical Decision Support Systems Based on Artificial Intelligence from Technological and Medical Perspective: A Systematic Review. [europepmc]
- Black-box assisted medical decisions: AI power vs. ethical physician care. [europepmc]
- Defending explicability as a principle for the ethics of artificial intelligence in medicine. [europepmc]
- Evaluating ChatGPT-4V in chest CT diagnostics: a critical image interpretation assessment. [europepmc]
- A systematic review of machine learning applications in predicting opioid associated adverse events. [europepmc]
Related