Consistent Individualized Feature Attribution for Tree Ensembles
2018/02/12 by Scott Lundberg, Scott M. Lundberg, Gabriel G. Erion +6 · 111 citations
Computer Science · Environmental Science · Mathematics · #Data Analysis with R #Explainable Artificial Intelligence (XAI) #Forest ecology and management #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1802.03888
Follow-up to 2017 ICML Workshop arXiv:1706.06060
arxiv created 2019/03/07 · arxiv updated 2019/03/08
Abstract
Interpreting predictions from tree ensemble methods such as gradient boosting machines and random forests is important, yet feature attribution for trees is often heuristic and not individualized for each prediction. Here we show that popular feature attribution methods are inconsistent, meaning they can lower a feature's assigned importance when the true impact of that feature actually increases. This is a fundamental problem that casts doubt on any comparison between features. To address it we turn to recent applications of game theory and develop fast exact tree solutions for SHAP (SHapley Additive exPlanation) values, which are the unique consistent and locally accurate attribution values. We then extend SHAP values to interaction effects and define SHAP interaction values. We propose a rich visualization of individualized feature attributions that improves over classic attribution summaries and partial dependence plots, and a unique "supervised" clustering (clustering based on feature attributions). We demonstrate better agreement with human intuition through a user study, exponential improvements in run time, improved clustering performance, and better identification of influential features. An implementation of our algorithm has also been merged into XGBoost and LightGBM, see http://github.com/slundberg/shap for details.
Citations
Cited by
- Multi-component Dark Matter in a Novel Three-Loop Inverse Scotogenic Seesaw Model
- Learning from sanctioned government suppliers: A machine learning and network science approach to detecting fraud and corruption in Mexico
- The Effect of Enforcing Fairness on Reshaping Explanations in Machine Learning Models
- Evaluating the Correctness of Explainable AI Algorithms for Classification
- Understanding peacefulness through the world news
- SX-GeoTree: Self-eXplaining Geospatial Regression Tree Incorporating the Spatial Similarity of Feature Attributions
- From Decision Trees to Boolean Logic: A Fast and Unified SHAP Algorithm
- Problems with Shapley-value-based explanations as feature importance measures
- Investigating U.S. Consumer Demand for Food Products with Innovative Transportation Certificates Based on Stated Preferences and Machine Learning Approaches
- Community Detection on Model Explanation Graphs for Explainable AI
- Discovering and Explaining the Representation Bottleneck of DNNs
- Explainable AI: current status and future directions
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainability
- Defining and Quantifying the Emergence of Sparse Concepts in DNNs
- On the Tractability of SHAP Explanations
- Bias, Fairness, and Accountability with AI and ML Algorithms
- The many Shapley values for model explanation
- Do Community Characteristics Explain Heat‐Related Illness in Seoul, Korea?
- Who stands out? Defining, measuring, and uncovering the elements of unusualness
- Feature Synergy, Redundancy, and Independence in Global Model Explanations using SHAP Vector Decomposition
- Tractable Shapley Values and Interactions via Tensor Networks
- The Need for Standardized Explainability
- An interpretable molecular descriptor for machine learning predictions in atmospheric science
- Data-Driven Analysis of Intersectional Bias in Image Classification: A Framework with Bias-Weighted Augmentation
- Clusters in Explanation Space: Inferring disease subtypes from model explanations
- Technical Note: Game-Theoretic Interactions of Different Orders
- SHAP-Based Supervised Clustering for Sample Classification and the Generalized Waterfall Plot
- Fourier Analysis on the Boolean Hypercube via Hoeffding Functional Decomposition
- Predictive Modeling and Explainable AI for Veterinary Safety Profiles, Residue Assessment, and Health Outcomes Using Real-World Data and Physicochemical Properties
- Adaptive Node Feature Selection For Graph Neural Networks
- A Human-Grounded Evaluation of SHAP for Alert Processing
- A Machine Learning Framework for Predicting and Understanding the Canadian Drought Monitor
- A weakened diurnal weather constraint leads to longer burning hours in North America
- The Role of Green Official Development Assistance in the Implementation of Sustainable Development Goal 15 Using Explainable AI
- Interpretable Anomaly Detection with DIFFI: Depth-based Isolation Forest Feature Importance
- Can We Faithfully Represent Masked States to Compute Shapley Values on a DNN?
- Enhancing ML Models Interpretability for Credit Scoring
- Neurosymbolic AI Transfer Learning Improves Network Intrusion Detection
- TREX: Tree-Ensemble Representer-Point Explanations
- XGBoostLSS -- An extension of XGBoost to probabilistic forecasting
- Combat COVID-19 Infodemic Using Explainable Natural Language Processing Models
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- Shapley Flow: A Graph-based Approach to Interpreting Model Predictions
- ELATE: Evolutionary Language model for Automated Time-series Engineering
- Shapley Values: Paired-Sampling Approximations
- Data-Driven Abdominal Phenotypes of Type 2 Diabetes in Lean, Overweight, and Obese Cohorts
- XGBoost meets INLA: a two-stage spatio-temporal forecasting of wildfires in Portugal
- High Dimensional Model Explanations: an Axiomatic Approach
- Holistic Explainable AI (H-XAI): Extending Transparency Beyond Developers in AI-Driven Decision Making
- Breaking New Ground in Software Defect Prediction: Introducing Practical and Actionable Metrics with Superior Predictive Power for Enhanced Decision-Making
- Case-Based Reasoning for Assisting Domain Experts in Processing Fraud Alerts of Black-Box Machine Learning Models
- An Unsupervised Deep Explainable AI Framework for Localization of Concurrent Replay Attacks in Nuclear Reactor Signals
- Generalized Integrated Gradients: A practical method for explaining diverse ensembles
- Who's responsible? Jointly quantifying the contribution of the learning algorithm and training data
- Incorporating Priors with Feature Attribution on Text Classification
- Do not explain without context: addressing the blind spot of model explanations
- Improved Feature Importance Computations for Tree Models: Shapley vs. Banzhaf
- Survival Regression with Accelerated Failure Time Model in XGBoost
- Explainable Evidential Clustering
- OrdShap: Feature Position Importance for Sequential Black-Box Models
- SHAFF: Fast and consistent SHApley eFfect estimates via random Forests
- Computational analysis of laminar structure of the human cortex based on local neuron features
- Feature relevance quantification in explainable AI: A causal problem
- Impact of Accuracy on Model Interpretations
- midr: Learning from Black-Box Models by Maximum Interpretation Decomposition
- Investigating and Simplifying Masking-based Saliency Methods for Model\n Interpretability
- Using Eye-tracking Data to Predict Situation Awareness in Real Time during Takeover Transitions in Conditionally Automated Driving
- From Street Form to Spatial Justice: Explaining Urban Exercise Inequality via a Triadic SHAP-Informed Framework
- Explaining deep neural network models for electricity price forecasting with XAI
- A Categorisation of Post-hoc Explanations for Predictive Models
- Learning to Ask Medical Questions using Reinforcement Learning
- Fast TreeSHAP: Accelerating SHAP Value Computation for Trees
- Regression-adjusted Monte Carlo Estimators for Shapley Values and Probabilistic Values
- Hyperparameter Optimization for Forecasting Stock Returns
- A Unified Framework for Provably Efficient Algorithms to Estimate Shapley Values
- Evolutionary model discovery of causal factors behind the socio-agricultural behavior of the ancestral Pueblo
- Inferring feature importance with uncertainties in high-dimensional data
- Neurosymbolic Artificial Intelligence for Robust Network Intrusion Detection: From Scratch to Transfer Learning
- Interpreting and Boosting Dropout from a Game-Theoretic View
- Shapley explainability on the data manifold
- Do Not Trust Additive Explanations
- Design Rules for Optimizing Quaternary Mixed-Metal Chalcohalides
- Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
- AI-Driven SEEG Channel Ranking for Epileptogenic Zone Localization
- Induction of Non-Monotonic Rules From Statistical Learning Models Using High-Utility Itemset Mining
- Relate and Predict: Structure-Aware Prediction with Jointly Optimized Neural DAG
- RePPL: Recalibrating Perplexity by Uncertainty in Semantic Propagation and Language Generation for Explainable QA Hallucination Detection
- Connecting Interpretability and Robustness in Decision Trees through Separation
- Surrogate Interpretable Graph for Random Decision Forests
- Importance measures derived from random forests: characterisation and extension
- Machine learning in the social and health sciences
- Financial Fraud Detection Using Explainable AI and Stacking Ensemble Methods
- On the Art and Science of Machine Learning Explanations
- Can I trust you more? Model-Agnostic Hierarchical Explanations
- Consistent feature attribution for tree ensembles
- A Debiased MDI Feature Importance Measure for Random Forests
- Modeling Dispositional and Initial learned Trust in Automated Vehicles with Predictability and Explainability
- DNN2LR: Automatic Feature Crossing for Credit Scoring
- Detecting and Explaining Unlawful Insider Trading: A Shapley Value and Causal Forest Approach to Identifying Key Drivers and Causal Relationships
- Beyond the Numbers: Causal Effects of Financial Report Sentiment on Bank Profitability
- Definitions, methods, and applications in interpretable machine learning
- What's Wrong with Your Synthetic Tabular Data? Using Explainable AI to Evaluate Generative Models
- Beyond the QBER Threshold: A Temporal QBER Based Machine Learning Framework for Multi Attack Detection in BB84 QKD
- Transferable and Extensible Machine Learning-Derived Atomic Charges for Modeling Hybrid Nanoporous Materials
- Scalable and Generalizable Social Bot Detection through Data Selection
- Interpreted machine learning in fluid dynamics: explaining relaminarisation events in wall-bounded shear flows
- WRSE -- a non-parametric weighted-resolution ensemble for predicting\n individual survival distributions in the ICU
- Explainable Artificial Intelligence for Process Mining: A General Overview and Application of a Novel Local Explanation Approach for Predictive Process Monitoring
- Nitrogen deposition favors later leaf senescence in woody species
- Explaining Anomalies Detected by Autoencoders Using SHAP
- Quantum Phases Classification Using Quantum Machine Learning with SHAP-Driven Feature Selection
- Regretful Decisions under Label Noise
- Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms
Related