Consistent Individualized Feature Attribution for Tree Ensembles
2018/02/12 by Scott Lundberg, Lundberg, Scott M., Gabriel Erion +3 · 63 citations
Environmental Science · Computer Science · #Forest ecology and management #Explainable Artificial Intelligence (XAI) #Data Analysis with R
paper · pdf · doi:10.48550/arxiv.1802.03888
Abstract
Interpreting predictions from tree ensemble methods such as gradient boosting machines and random forests is important, yet feature attribution for trees is often heuristic and not individualized for each prediction. Here we show that popular feature attribution methods are inconsistent, meaning they can lower a feature's assigned importance when the true impact of that feature actually increases. This is a fundamental problem that casts doubt on any comparison between features. To address it we turn to recent applications of game theory and develop fast exact tree solutions for SHAP (SHapley Additive exPlanation) values, which are the unique consistent and locally accurate attribution values. We then extend SHAP values to interaction effects and define SHAP interaction values. We propose a rich visualization of individualized feature attributions that improves over classic attribution summaries and partial dependence plots, and a unique "supervised" clustering (clustering based on feature attributions). We demonstrate better agreement with human intuition through a user study, exponential improvements in run time, improved clustering performance, and better identification of influential features. An implementation of our algorithm has also been merged into XGBoost and LightGBM, see http://github.com/slundberg/shap for details.
Cited by
- Multi-component Dark Matter in a Novel Three-Loop Inverse Scotogenic Seesaw Model
- Learning from sanctioned government suppliers: A machine learning and network science approach to detecting fraud and corruption in Mexico
- The Effect of Enforcing Fairness on Reshaping Explanations in Machine Learning Models
- Evaluating the Correctness of Explainable AI Algorithms for Classification
- Understanding peacefulness through the world news
- SX-GeoTree: Self-eXplaining Geospatial Regression Tree Incorporating the Spatial Similarity of Feature Attributions
- From Decision Trees to Boolean Logic: A Fast and Unified SHAP Algorithm
- Problems with Shapley-value-based explanations as feature importance\n measures
- Investigating U.S. Consumer Demand for Food Products with Innovative Transportation Certificates Based on Stated Preferences and Machine Learning Approaches
- Community Detection on Model Explanation Graphs for Explainable AI
- Discovering and Explaining the Representation Bottleneck of DNNs
- Explainable AI: current status and future directions
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainability
- Defining and Quantifying the Emergence of Sparse Concepts in DNNs
- On the Tractability of SHAP Explanations
- Bias, Fairness, and Accountability with AI and ML Algorithms
- The many Shapley values for model explanation
- Do Community Characteristics Explain Heat‐Related Illness in Seoul, Korea?
- Who stands out? Defining, measuring, and uncovering the elements of unusualness
- Feature Synergy, Redundancy, and Independence in Global Model Explanations using SHAP Vector Decomposition
- Tractable Shapley Values and Interactions via Tensor Networks
- The Need for Standardized Explainability
- An interpretable molecular descriptor for machine learning predictions in atmospheric science
- Data-Driven Analysis of Intersectional Bias in Image Classification: A Framework with Bias-Weighted Augmentation
- Clusters in Explanation Space: Inferring disease subtypes from model explanations
- Technical Note: Game-Theoretic Interactions of Different Orders
- SHAP-Based Supervised Clustering for Sample Classification and the Generalized Waterfall Plot
- Fourier Analysis on the Boolean Hypercube via Hoeffding Functional Decomposition
- Predictive Modeling and Explainable AI for Veterinary Safety Profiles, Residue Assessment, and Health Outcomes Using Real-World Data and Physicochemical Properties
- Adaptive Node Feature Selection For Graph Neural Networks
- A Human-Grounded Evaluation of SHAP for Alert Processing
- A Machine Learning Framework for Predicting and Understanding the Canadian Drought Monitor
- A weakened diurnal weather constraint leads to longer burning hours in North America
- The Role of Green Official Development Assistance in the Implementation of Sustainable Development Goal 15 Using Explainable AI
- Interpretable Anomaly Detection with DIFFI: Depth-based Isolation Forest Feature Importance
- Can We Faithfully Represent Masked States to Compute Shapley Values on a DNN?
- Enhancing ML Models Interpretability for Credit Scoring
- Neurosymbolic AI Transfer Learning Improves Network Intrusion Detection
- TREX: Tree-Ensemble Representer-Point Explanations
- XGBoostLSS -- An extension of XGBoost to probabilistic forecasting
- Combat COVID-19 Infodemic Using Explainable Natural Language Processing Models
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- Shapley Flow: A Graph-based Approach to Interpreting Model Predictions
- ELATE: Evolutionary Language model for Automated Time-series Engineering
- Shapley Values: Paired-Sampling Approximations
- Data-Driven Abdominal Phenotypes of Type 2 Diabetes in Lean, Overweight, and Obese Cohorts
- XGBoost meets INLA: a two-stage spatio-temporal forecasting of wildfires in Portugal
- High Dimensional Model Explanations: an Axiomatic Approach
- Holistic Explainable AI (H-XAI): Extending Transparency Beyond Developers in AI-Driven Decision Making
- Breaking New Ground in Software Defect Prediction: Introducing Practical and Actionable Metrics with Superior Predictive Power for Enhanced Decision-Making
- Case-Based Reasoning for Assisting Domain Experts in Processing Fraud Alerts of Black-Box Machine Learning Models
- An Unsupervised Deep Explainable AI Framework for Localization of Concurrent Replay Attacks in Nuclear Reactor Signals
- Generalized Integrated Gradients: A practical method for explaining diverse ensembles
- Who's responsible? Jointly quantifying the contribution of the learning algorithm and training data
- Incorporating Priors with Feature Attribution on Text Classification
- Do not explain without context: addressing the blind spot of model explanations
- Improved Feature Importance Computations for Tree Models: Shapley vs. Banzhaf
- Survival Regression with Accelerated Failure Time Model in XGBoost
- Explainable Evidential Clustering
- OrdShap: Feature Position Importance for Sequential Black-Box Models
- SHAFF: Fast and consistent SHApley eFfect estimates via random Forests
- Computational analysis of laminar structure of the human cortex based on local neuron features
- Feature relevance quantification in explainable AI: A causal problem
Related