"Why Should I Trust You?": Explaining the Predictions of Any Classifier
2016/02/16 by Marco Túlio Ribeiro, Marco Tulio Ribeiro, Ribeiro, Marco Tulio +4 · 5 voices · 451 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Machine Learning and Data Classification #cs.AI #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1602.04938
arxiv published 2016/02/16 · arxiv updated 2016/08/09
Abstract
Despite widespread adoption, machine learning models remain mostly black boxes. Understanding the reasons behind predictions is, however, quite important in assessing trust, which is fundamental if one plans to take action based on a prediction, or when choosing whether to deploy a new model. Such understanding also provides insights into the model, which can be used to transform an untrustworthy model or prediction into a trustworthy one. In this work, we propose LIME, a novel explanation technique that explains the predictions of any classifier in an interpretable and faithful manner, by learning an interpretable model locally around the prediction. We also propose a method to explain models by presenting representative individual predictions and their explanations in a non-redundant way, framing the task as a submodular optimization problem. We demonstrate the flexibility of these methods by explaining different models for text (e.g. random forests) and image classification (e.g. neural networks). We show the utility of explanations via novel experiments, both simulated and with human subjects, on various scenarios that require trust: deciding if one should trust a prediction, choosing between models, improving an untrustworthy classifier, and identifying why a classifier should not be trusted.
Citations
Cited by
- Explainable Neural Inverse Kinematics for Obstacle-Aware Robotic Manipulation: A Comparative Analysis of IKNet Variants
- MiSS: A Logic-Driven Explanation of Minimal Sufficient Coalitions for Point Cloud Classifiers
- A CNN-Based Malaria Diagnosis from Blood Cell Images with SHAP and LIME Explainability
- Explainable Model Routing for Agentic Workflows
- LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
- Context-Aware Concept Distillation for Trustworthy Flood Prediction
- Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
- From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation
- Explainable Multimodal Regression via Information Decomposition
- Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations
- Beyond Adoption Intention How Trust in Augmented Analytics Relates to Perceived Decision Quality Among Non-Technical BI Users
- A Unified Framework for Uncertainty-Aware Explainable Artificial Intelligence: A Case Study in Power Quality Disturbance Classification
- Explainability Methods for Hardware Trojan Detection: A Systematic Comparison
- A Model of Causal Explanation on Neural Networks for Tabular Data
- Towards Generative Location Awareness for Disaster Response: A Probabilistic Cross-view Geolocalization Approach
- Explainable time-series forecasting with sampling-free SHAP for Transformers
- Toward Explaining Large Language Models in Software Engineering Tasks
- UbiQVision: Quantifying Uncertainty in XAI for Image Recognition
- Reason2Decide: Rationale-Driven Multi-Task Learning
- DecoKAN: Interpretable Decomposition for Forecasting Cryptocurrency Market Dynamics
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Augmenting Intelligence: A Hybrid Framework for Scalable and Stable Explanations
- Time-series Forecast for Indoor Zone Air Temperature with Long Horizons: A Case Study with Sensor-based Data from a Smart Building
- Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
- PROVEX: Enhancing SOC Analyst Trust with Explainable Provenance-Based IDS
- Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
- XAgen: An Explainability Tool for Identifying and Correcting Failures in Multi-Agent Workflows
- Interpretable Plant Leaf Disease Detection Using Attention-Enhanced CNN
- Adversarially Robust Detection of Harmful Online Content: A Computational Design Science Approach
- Explanation Beyond Intuition: A Testable Criterion for Inherent Explainability
- Enhancing Long Document Long Form Summarisation with Self-Planning
- UniCoMTE: A Universal Counterfactual Framework for Explaining Time-Series Classifiers on ECG Data
- Metanetworks as Regulatory Operators: Learning to Edit for Requirement Compliance
- Explainable AI in Big Data Fraud Detection
- Explaining the Reasoning of Large Language Models Using Attribution Graphs
- Multi-Agent Medical Decision Consensus Matrix System: An Intelligent Collaborative Framework for Oncology MDT Consultations
- Enhancing Interpretability for Vision Models via Shapley Value Optimization
- Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Stud
- Explainable reinforcement learning from human feedback to improve alignment
- Interpretable Hypothesis-Driven Trading:A Rigorous Walk-Forward Validation Framework for Market Microstructure Signals
- AgentSHAP: Interpreting LLM Agent Tool Importance with Monte Carlo Shapley Value Estimation
- Explainable Prediction of Economic Time Series Using IMFs and Neural Networks
- SRLR: Symbolic Regression based Logic Recovery to Counter Programmable Logic Controller Attacks
- Back to the Baseline: Examining Baseline Effects on Explainability Metrics
- Learning complete and explainable visual representations from itemized text supervision
- Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution
- Provable Recovery of Locally Important Signed Features and Interactions from Random Forest
- How Interpretable and Trustworthy are GAMs?
- Interpreto: An Explainability Library for Transformers
- Explainable Verification of Hierarchical Workflows Mined from Event Logs with Shapley Values
- TRUCE: TRUsted Compliance Enforcement Service for Secure Health Data Exchange
- SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning
- Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts
- Hybrid Attribution Priors for Explainable and Robust Model Training
- SSplain: Sparse and Smooth Explainer for Retinopathy of Prematurity Classification
- HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability
- ϕ-Table: A Statistical Explanation for Global SHAP
- A Geometric Unification of Concept Learning with Concept Cones
- CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
- Always Keep Your Promises: A Model-Agnostic Attribution Algorithm for Neural Networks
- CLUENet: Cluster Attention Makes Neural Networks Have Eyes
- Automated Plant Disease and Pest Detection System Using Hybrid Lightweight CNN-MobileViT Models for Diagnosis of Indigenous Crops
- MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution
- Improving Local Fidelity Through Sampling and Modeling Nonlinearity
- Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches
- Lyrics Matter: Exploiting the Power of Learnt Representations for Music Popularity Prediction
- Sell Data to AI Algorithms Without Revealing It: Secure Data Valuation and Sharing via Homomorphic Encryption
- Sample-Efficient Counterfactual Tuning for Compressor Pressure Control
- MASE: Interpretable NLP Models via Model-Agnostic Saliency Estimation
- Explainable AI for Smart Greenhouse Control: Interpretability of Temporal Fusion Transformer in the Internet of Robotic Things
- Out-of-the-box: Black-box Causal Attacks on Object Detectors
- A Framework for Causal Concept-based Model Explanations
- Beyond Additivity: Sparse Isotonic Shapley Regression toward Nonlinear Explainability
- Training Data Attribution for Image Generation using Ontology-Aligned Knowledge Graphs
- Demystifying Feature Engineering in Malware Analysis of API Call Sequences
- Forecasting India's Demographic Transition Under Fertility Policy Scenarios Using hybrid LSTM-PINN Model
- Supervised Contrastive Machine Unlearning of Background Bias in Sonar Image Classification with Fine-Grained Explainable AI
- Building Trustworthy AI for Materials Discovery: From Autonomous Laboratories to Z-scores
- Pushing the Boundaries of Interpretability: Incremental Enhancements to the Explainable Boosting Machine
- Neural Networks as Physics-Consistent Surrogates: An Explainable AI Validation Framework for Learning Constitutive Relations
- CodeFlowLM: Incremental Just-In-Time Defect Prediction with Pretrained Language Models and Exploratory Insights into Defect Localization
- From 'What-is' to 'What-if' in Human-Factor Analysis: A Post-Occupancy Evaluation Case
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model Predictions
- Faster Verified Explanations for Neural Networks
- DeFi TrustBoost: Blockchain and AI for Trustworthy Decentralized Financial Decisions
- Does the Model Say What the Data Says? A Simple Heuristic for Model Data Alignment
- Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring
- Analysis of Invasive Breast Cancer in Mammograms Using YOLO, Explainability, and Domain Adaptation
- BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning
- On Rank Graduation Metrics for High Dimensional Ordinal Data
- SHIC-XE: Viewpoint-Invariant Explainability via Dense 2D-3D Correspondences: an Application to Equine Pain Recognition
- Space Explanations of Neural Network Classification
- MATCH: Engineering Transparent and Controllable Conversational XAI Systems through Composable Building Blocks
- Physics-Informed Spiking Neural Networks via Conservative Flux Quantization
- The Directed Prediction Change - Efficient and Trustworthy Fidelity Assessment for Local Feature Attribution Methods
- Accumulated Local Effects and Graph Neural Networks for link prediction
- Ranking-Enhanced Anomaly Detection Using Active Learning-Assisted Attention Adversarial Dual AutoEncoders
- Generation, Evaluation, and Explanation of Novelists' Styles with Single-Token Prompts
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- Think Fast: Real-Time IoT Intrusion Reasoning Using IDS and LLMs at the Edge Gateway
- Matching-Based Few-Shot Semantic Segmentation Models Are Interpretable by Design
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
- Correlation-Aware Feature Attribution Based Explainable AI
- CID: Measuring Feature Importance Through Counterfactual Distributions
- MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm
- From Black Box to Insight: Explainable AI for Extreme Event Preparedness
- Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
- Multi-task GINN-LP for Multi-target Symbolic Regression
- From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability
- SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data
- LAYA: Layer-wise Attention Aggregation for Interpretable Depth-Aware Neural Networks
- Explainable Recommender Systems via Resolving Learning Representations
- Interpretable Fine-Gray Deep Survival Model for Competing Risks: Predicting Post-Discharge Foot Complications for Diabetic Patients in Ontario
- SCI: A Metacognitive Control for Signal Dynamics
- A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
- Additive Large Language Models for Semi-Structured Text
- ClinStructor: AI-Powered Structuring of Unstructured Clinical Texts
- Multi-omic Enriched Blood-Derived Digital Signatures Reveal Mechanistic and Confounding Disease Clusters for Differential Diagnosis
- Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
- Privacy-Preserving Explainable AIoT Application via SHAP Entropy Regularization
- When Thinking Pays Off: Incentive Alignment for Human-AI Collaboration
- MACIE: Multi-Agent Causal Intelligence Explainer for Collective Behavior Understanding
- llmSHAP: A Principled Approach to LLM Explainability
- Contrastive Integrated Gradients: A Feature Attribution-Based Method for Explaining Whole Slide Image Classification
- Improving Industrial Injection Molding Processes with Explainable AI for Quality Classification
- Increasing AI Explainability by LLM Driven Standard Processes
- Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
- Design Principles of Zero-Shot Self-Supervised Unknown Emitter Detectors
- Explainable Probabilistic Machine Learning for Predicting Drilling Fluid Loss of Circulation in Marun Oil Field
- Enhancing Adversarial Robustness of IoT Intrusion Detection via SHAP-Based Attribution Fingerprinting
- Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks
- DiagnoLLM: A Hybrid Bayesian Neural Language Framework for Interpretable Disease Diagnosis
- Adversarially Robust and Interpretable Magecart Malware Detection
- Unlocking the Black Box: A Five-Dimensional Framework for Evaluating Explainable AI in Credit Risk
- Uncovering Bugs in Formal Explainers: A Case Study with PyXAI
- Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
- Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
- Fair and Explainable Credit-Scoring under Concept Drift: Adaptive Explanation Frameworks for Evolving Populations
- Connecting the concepts of quantum state tomography and molecular representations for machine learning
- Geometry as a Missing Axis of Representation Quality: The Variational Geometric Information Bottleneck under Data Scarcity
- LLEXICORP: End-user Explainability of Convolutional Neural Networks
- AgentSLA : Towards a Service Level Agreement for AI Agents
- Community Detection on Model Explanation Graphs for Explainable AI
- Imbalanced Classification through the Lens of Spurious Correlations
- Multilingual BERT language model for medical tasks: Evaluation on domain-specific adaptation and cross-linguality
- Interpretable Model-Aware Counterfactual Explanations for Random Forest
- Before the Clinic: Transparent and Operable Design Principles for Healthcare AI
- Feature-Function Curvature Analysis: A Geometric Framework for Explaining Differentiable Models
- Counterfactual-based Agent Influence Ranker for Agentic AI Workflows
- Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
- Explainable AI: current status and future directions
- XRAI: Better Attributions Through Regions
- RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
- (EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations
- Comparing Rule-Based and Deep Learning Models for Patient Phenotyping
- On a Method to Measure Supervised Multiclass Model’s Interpretability: Application to Degradation Diagnosis (Short Paper)
- Reward Models are Metrics in a Trench Coat
- Convolutional neural networks in Vis–NIR chemometrics: From contradiction to conditional design
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- Deep representation learning for temporal inference in cancer omics: a systematic literature review
- Towards Robust Interpretability with Self-Explaining Neural Networks
- The many Shapley values for model explanation
- Neurodatascience: Past, Present, and Future
- Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
- Dating Apps and the Right to an Explanation
- Strategic inputs: feature selection from game-theoretic perspective
- FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning
- A Novel XAI-Enhanced Quantum Adversarial Networks for Velocity Dispersion Modeling in MaNGA Galaxies
- Feature-Guided Analysis of Neural Networks: A Replication Study
- AI in soil moisture remote sensing
- Toward Carbon-Neutral Human AI: Rethinking Data, Computation, and Learning Paradigms for Sustainable Intelligence
- MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
- Rule-Based Explanations for Retrieval-Augmented LLM Systems
- Towards Piece-by-Piece Explanations for Chess Positions with SHAP
- Tractable Shapley Values and Interactions via Tensor Networks
- From Black-box to Causal-box: Towards Building More Interpretable Models
- Additive Models Explained: A Computational Complexity Approach
- Embedding Explainable AI in NHS Clinical Safety: The Explainability-Enabled Clinical Safety Framework (ECSF)
- Knowledge-Driven Vision-Language Model for Plexus Detection in Hirschsprung's Disease
- Framework for Machine Evaluation of Reasoning Completeness in Large Language Models For Classification Tasks
- Citation Failure: Definition, Analysis and Efficient Mitigation
- Human-Centered LLM-Agent System for Detecting Anomalous Digital Asset Transactions
- ShapeX: Shapelet-Driven Post Hoc Explanations for Time Series Classification Models
- Green Finance and Carbon Emissions: A Nonlinear and Interaction Analysis Using Bayesian Additive Regression Trees
- From Prototypes to Sparse ECG Explanations: SHAP-Driven Counterfactuals for Multivariate Time-Series Multi-class Classification
- Inference on Variable Importance for Treatment Effect Heterogeneity: Shapley Values and Beyond
- Motion2Meaning: A Clinician-Centered Framework for Contestable LLM in Parkinson's Disease Gait Interpretation
- DSEBench: A Test Collection for Explainable Dataset Search with Examples
- Preliminary Quantitative Study on Explainability and Trust in AI Systems
- Reflections from Research Roundtables at the Conference on Health, Inference, and Learning (CHIL) 2025
- LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching
- Learnable Game-theoretic Policy Optimization for Data-centric Self-explanation Rationalization
- XD-RCDepth: Lightweight Radar-Camera Depth Estimation with Explainability-Aligned and Distribution-Aware Distillation
- Doctor Rashomon and the UNIVERSE of Madness: Variable Importance with Unobserved Confounding and the Rashomon Effect
- ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
- Unveiling the Vulnerability of Graph-LLMs: An Interpretable Multi-Dimensional Adversarial Attack on TAGs
- On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- Beyond single-model XAI: aggregating multi-model explanations for enhanced trustworthiness
- A Comprehensive Forecasting-Based Framework for Time Series Anomaly Detection: Benchmarking on the Numenta Anomaly Benchmark (NAB)
- Evaluating the Explainability of Vision Transformers in Medical Imaging
- Restricted Receptive Fields for Face Verification
- Assessing Policy Updates: Toward Trust-Preserving Intelligent User Interfaces
- f-INE: A Hypothesis Testing Framework for Estimating Influence under Training Randomness
- Denoised IPW-Lasso for Heterogeneous Treatment Effect Estimation in Randomized Experiments
- SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
- Explainable Human-in-the-Loop Segmentation via Critic Feedback Signals
- The Ethics Engine: A Modular Pipeline for Accessible Psychometric Assessment of Large Language Models
- Data Science in the Big Data Era: Analytics, Intelligence, and Future Challenges
- The Evolution of Artificial Intelligence Paradigms – Implications for Performance, Scalability, and Responsible Deployment
- Interpretable Generative and Discriminative Learning for Multimodal and Incomplete Clinical Data
- Towards Safer and Understandable Driver Intention Prediction
- It's 2025 -- Narrative Learning is the new baseline to beat for explainable machine learning
- From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening
- CARLE: A Hybrid Deep-Shallow Learning Framework for Robust and Explainable RUL Estimation of Rolling Element Bearings
- Training Feature Attribution for Vision Models
- Chain-of-Influence: Tracing Interdependencies Across Time and Features in Clinical Predictive Modelings
- Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
- Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
- Surrogate Graph Partitioning for Spatial Prediction
- IKNet: Interpretable Stock Price Prediction via Keyword-Guided Integration of News and Technical Indicators
- Machine Learning Algorithms for Financial Asset Price Forecasting
- Development of Mental Models in Human-AI Collaboration: A Conceptual Framework
- Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
- Enhancing Concept Localization in CLIP-based Concept Bottleneck Models
- Higher-Order Feature Attribution: Bridging Statistics, Explainable AI, and Topological Signal Processing
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
- The Role of Federated Learning in Improving Financial Security: A Survey
- QGraphLIME - Explaining Quantum Graph Neural Networks
- XAI-on-RAN: Explainable, AI-native, and GPU-Accelerated RAN Towards 6G
- Design Process of a Self Adaptive Smart Serious Games Ecosystem
- Reconsidering Requirements Engineering: Human-AI Collaboration in AI-Native Software Development
- Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
- Explainable but Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications
- Adaptive Node Feature Selection For Graph Neural Networks
- Onto-Epistemological Analysis of AI Explanations
- Provenance Networks: End-to-End Exemplar-Based Explainability
- Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
- Batch-CAM: Introduction to better reasoning in convolutional deep learning models
- Bayesian Neural Networks for Functional ANOVA model
- GDLNN: Marriage of Programming Language and Neural Networks for Accurate and Easy-to-Explain Graph Classification
- Explainable machine learning for multiscale thermal conductivity modeling in polymer nanocomposites with uncertainty quantification
- o-MEGA: Optimized Methods for Explanation Generation and Analysis
- Object-Centric Case-Based Reasoning via Argumentation
- MUSE-Explainer: Counterfactual Explanations for Symbolic Music Graph Classification Models
- ACT: Agentic Classification Tree
- ACE: Adapting sampling for Counterfactual Explanations
- Elucidating the Grey Atmosphere: SHAP Value Analysis of a Random Forest Atmospheric Neutral Density Model
- An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
- DeepProv: Behavioral Characterization and Repair of Neural Networks via Inference Provenance Graph Analysis
- Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- A Fuzzy Logic-Based Framework for Explainable Machine Learning in Big Data Analytics
- Identifying Information-Transfer Nodes in a Recurrent Neural Network Reveals Dynamic Representations
- Understanding Collaboration between Professional Designers and Decision-making AI: A Case Study in the Workplace
- Fidelity-Aware Data Composition for Robust Robot Generalization
- Fast and accurate explanations of distance-based classifiers by uncovering latent explanatory structures
- Improving Simple Models with Confidence Profiles
- Deep learning-based morphological classification of ceramics: A case study of 3D point cloud analysis for Sue ware, Japan
- Who Trusts AI for Health Information? A Cross-National Survey on Trust Determinants in Four European Countries
- On The Variability of Concept Activation Vectors
- A Novel Hybrid Deep Learning and Chaotic Dynamics Approach for Thyroid Cancer Classification
- CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
- SHAPoint: Task-Agnostic, Efficient, and Interpretable Point-Based Risk Scoring via Shapley Values
- Accuracy-Robustness Trade Off via Spiking Neural Network Gradient Sparsity Trail
- ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series Forecasting
- Towards Human-interpretable Explanation in Code Clone Detection using LLM-based Post Hoc Explainer
- MindCraft: How Concept Trees Take Shape In Deep Models
- Learning Temporal Saliency for Time Series Forecasting with Cross-Scale Attention
- CausalKANs: interpretable treatment effect estimation with Kolmogorov-Arnold networks
- Explaining multimodal LLMs via intra-modal token interactions
- Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
- Bridging the Gap Between Scientific Laws Derived by AI Systems and Canonical Knowledge via Abductive Inference with AI-Noether
- Optimal Robust Recourse with Lp-Bounded Model Change
- Explaining Fine Tuned LLMs via Counterfactuals A Knowledge Graph Driven Framework
- Learning Conformal Explainers for Image Classifiers
- EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email Defense
- GenFacts-Generative Counterfactual Explanations for Multi-Variate Time Series
- ExpIDS: A Drift-adaptable Network Intrusion Detection System With Improved Explainability
- Investigating Modality Contribution in Audio LLMs for Music
- Exact Subgraph Isomorphism Network with Mixed L0,2 Norm Constraint for Predictive Graph Mining
- TSKAN: Interpretable Machine Learning for QoE modeling over Time Series Data
- Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning
- Every Character Counts: From Vulnerability to Defense in Phishing Detection
- Trust and Transparency in AI: Industry Voices on Data, Ethics, and Compliance
- HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data
- Improving Mental Health Screening and Early Risk Detection in Spanish
- Class-Aware Reinforcement Learning for Counterfactual Explanation Generation
- FADEx: Feature Attribution and Distortion-based Explanation of Dimensionality Reduction
- Quantitative Evaluations on Saliency Methods: An Experimental Study
- Explainable AI Isn't Enough! Rethinking Algorithmic Contestability
- Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments
- Physics-constrained neural ordinary differential equation models to discover and predict microbial community dynamics
- Concerning Uncertainty -- A Systematic Survey of Uncertainty-Aware XAI
- Towards Transparent Application of Machine Learning in Video Processing
- FA(IR)2MA-GLVQ – A hidden-feature-bias mitigation approach for fairness in classification learning based on generalized matrix learning vector quantization
- Toward Human-Centered Explainability: Natural Language Explanations for Anomaly Detection
- Artificial Intelligence-Powered Raman Spectroscopy through Open Science and FAIR Principles
- Explaining and interpreting hyperdimensional computing classifiers on tabular data
- Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
- Learning From Simulators: A Theory of Simulation-Grounded Learning
- Path-Weighted Integrated Gradients for Interpretable Dementia Classification
- Personalized Prediction By Learning Halfspace Reference Classes Under Well-Behaved Distribution
- Building Transparency in Deep Learning-Powered Network Traffic Classification: A Traffic-Explainer Framework
- Attention Consistency for LLMs Explanation
- Looking in the mirror: A faithful counterfactual explanation method for interpreting deep image classification models
- Towards a Transparent and Interpretable AI Model for Medical Image Classifications
- Interpretable Clinical Classification with Kolmogorov-Arnold Networks
- CausalSent: Interpretable Sentiment Classification with RieszNet
- Benchmarking Class Activation Map Methods for Explainable Brain Hemorrhage Classification on Hemorica Dataset
- Learning Interpretable Differentiable Logic Networks for Time-Series Classification
- RulER: Automated Rule-Based Semantic Error Localization and Repair for Code Translation
- Copycat vs. Original: Multi-modal Pretraining and Variable Importance in Box-office Prediction
- NeRF-based Visualization of 3D Cues Supporting Data-Driven Spacecraft Pose Estimation
- Attention Lattice Adapter: Visual Explanation Generation for Visual Foundation Model
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- LLM on a Budget: Active Knowledge Distillation for Efficient Classification of Large Text Corpora
- Out of Distribution Detection in Self-adaptive Robots with AI-powered Digital Twins
- L-XAIDS: A LIME-based eXplainable AI framework for Intrusion Detection Systems
- Explainable Counterfactual Reasoning in Depression Medication Selection at Multi-Levels (Personalized and Population)
- An empirical comparison of machine learning methods for text-based sentiment analysis of online consumer reviews
- Regulating genome language models: navigating policy challenges at the intersection of AI and genetics
- Evalet: Evaluating Large Language Models through Functional Fragmentation
- Clarifying Model Transparency: Interpretability versus Explainability in Deep Learning with MNIST and IMDB Examples
- CUBE: Contrastive Understanding by Balanced Experiments
- Progressive Disclosure: Designing for Effective Transparency
- NuGraph2 with Explainability: Post-hoc Explanations for Geometric Neural Network Predictions
- Can Explainable AI Explain Unfairness? A Framework for Evaluating Explainable AI
- Model-agnostic post-hoc explainability for recommender systems
- Investigating Feature Attribution for 5G Network Intrusion Detection
- Opening the Black Box: Interpretable LLMs via Semantic Resonance Architecture
- Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
- Evaluation of Black-Box XAI Approaches for Predictors of Values of Boolean Formulae
- An Autoencoder and Vision Transformer-based Interpretability Analysis of the Differences in Automated Staging of Second and Third Molars
- Database Views as Explanations for Relational Deep Learning
- STRIDE: Subset-Free Functional Decomposition for XAI in Tabular Settings
- Uncertainty Awareness and Trust in Explainable AI- On Trust Calibration using Local and Global Explanations
- Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"
- An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
- Explainability of CNN Based Classification Models for Acoustic Signal
- Minimal Data, Maximum Clarity: A Heuristic for Explaining Optimization
- An Interpretable Deep Learning Model for General Insurance Pricing
- Transparency of medical artificial intelligence systems
- Hybrid GCN-GRU Model for Anomaly Detection in Cryptocurrency Transactions
- MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
- Expert-Guided Explainable Few-Shot Learning for Medical Image Diagnosis
- Rule-Based Moral Principles for Explaining Uncertainty in Natural Language Generation
- An Explainable Deep Neural Network with Frequency-Aware Channel and Spatial Refinement for Flood Prediction in Sustainable Cities
- A Survey of Challenges and Opportunities in Sensing and Analytics for\n Cardiovascular Disorders
- An Approach to Grounding AI Model Evaluations in Human-derived Criteria
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- BIDO: A Unified Approach to Address Obfuscation and Concept Drift Challenges in Image-based Malware Detection
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- Systematic Evaluation of Attribution Methods: Eliminating Threshold Bias and Revealing Method-Dependent Performance Patterns
- Enhancing Interpretability and Effectiveness in Recommendation with Numerical Features via Learning to Contrast the Counterfactual samples
- A XAI-based Framework for Frequency Subband Characterization of Cough Spectrograms in Chronic Respiratory Disease
- Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring
- Who Owns The Robot?: Four Ethical and Socio-technical Questions about Wellbeing Robots in the Real World through Community Engagement
- An Information-Flow Perspective on Explainability Requirements: Specification and Verification
- Wrong Model, Right Uncertainty: Spatial Associations for Discrete Data with Misspecification
- DRetNet: A Novel Deep Learning Framework for Diabetic Retinopathy Diagnosis
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- An Explainable Gaussian Process Auto-encoder for Tabular Data
- Causal SHAP: Feature Attribution with Dependency Awareness through Causal Discovery
- Tabular Diffusion Counterfactual Explanations
- XAI-Driven Machine Learning System for Driving Style Recognition and Personalized Recommendations
- Large Language Model Integration with Reinforcement Learning to Augment Decision-Making in Autonomous Cyber Operations
- GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability
- Individualized and Interpretable Sleep Forecasting via a Two-Stage Adaptive Spatial-Temporal Model
- Axiomatic Attribution for Deep Networks
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating
- Multi-View Graph Convolution Network for Internal Talent Recommendation Based on Enterprise Emails
- Automated Quality Assessment for LLM-Based Complex Qualitative Coding: A Confidence-Diversity Framework
- Safer Skin Lesion Classification with Global Class Activation Probability Map Evaluation and SafeML
- A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
- Attention-Based Explainability for Structure-Property Relationships
- Breaking the Black Box: Inherently Interpretable Physics-Constrained Machine Learning With Weighted Mixed-Effects for Imbalanced Seismic Data
- VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- PAX-TS: Model-agnostic multi-granular explanations for time series forecasting via localized perturbations
- How Reliable are LLMs for Reasoning on the Re-ranking task?
- Locally Pareto-Optimal Interpretations for Black-Box Machine Learning Models
- The Loupe: A Plug-and-Play Attention Module for Amplifying Discriminative Features in Vision Transformers
- QA-VLM: Providing human-interpretable quality assessment for wire-feed laser additive manufacturing parts with Vision Language Models
- XAI-Driven Spectral Analysis of Cough Sounds for Respiratory Disease Characterization
- Exact Shapley Attributions in Quadratic-time for FANOVA Gaussian Processes
- A Comparative Evaluation of Teacher-Guided Reinforcement Learning Techniques for Autonomous Cyber Operations
- Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
- ITL-LIME: Instance-Based Transfer Learning for Enhancing Local Explanations in Low-Resource Data Settings
- Compressed Models are NOT Trust-equivalent to Their Large Counterparts
- Interpreting Time Series Forecasts with LIME and SHAP: A Case Study on the Air Passengers Dataset
- Design and Validation of a Responsible Artificial Intelligence-based System for the Referral of Diabetic Retinopathy Patients
- Rigorous Feature Importance Scores based on Shapley Value and Banzhaf Index
- Informative Post-Hoc Explanations Only Exist for Simple Functions
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- When Explainability Meets Privacy: An Investigation at the Intersection of Post-hoc Explainability and Differential Privacy in the Context of Natural Language Processing
- RealAC: A Domain-Agnostic Framework for Realistic and Actionable Counterfactual Explanations
- An unexpected unity among methods for interpreting model predictions
- Explainable AI Technique in Lung Cancer Detection Using Convolutional Neural Networks
- Extending the Entropic Potential of Events for Uncertainty Quantification and Decision-Making in Artificial Intelligence
- Adoption of Explainable Natural Language Processing: Perspectives from Industry and Academia on Practices and Challenges
- A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- Understanding Dementia Speech Alignment with Diffusion-Based Image Generation
- SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling
- Reveal-Bangla: A Dataset for Cross-Lingual Multi-Step Reasoning Evaluation
- Beyond Technocratic XAI: The Who, What & How in Explanation Design
- Hierarchical Variable Importance with Statistical Control for Medical Data-Based Prediction
- Beyond Predictions: A Study of AI Strength and Weakness Transparency Communication on Human-AI Collaboration
- Deep Reinforcement Learning with Local Interpretability for Transparent Microgrid Resilience Energy Management
- Exploring Content and Social Connections of Fake News with Explainable Text and Graph Learning
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
- Feature contributions and predictive modeling of aeolian sand transport detection in the atmospheric surface layer
- An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal
- AttnTrace: Attention-based Context Traceback for Long-Context LLMs
- SHAP Stability in Credit Risk Management: A Case Study in Credit Card Default Model
- TeSent: A Benchmark Dataset for Fairness-aware Explainable Sentiment Classification in Telugu
- AutoSIGHT: Automatic Eye Tracking-based System for Immediate Grading of Human experTise
- TofuML: A Spatio-Physical Interactive Machine Learning Device for Interactive Exploration of Machine Learning for Novices
- MetaExplainer: A Framework to Generate Multi-Type User-Centered Explanations for AI Systems
- Causal Identification of Sufficient, Contrastive and Complete Feature Sets in Image Classification
- Your Model Is Unfair, Are You Even Aware? Inverse Relationship Between Comprehension and Trust in Explainability Visualizations of Biased ML Models
- TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning
- Teaching the Teacher: Improving Neural Network Distillability for Symbolic Regression via Jacobian Regularization
- XAI for Point Cloud Data using Perturbations based on Meaningful Segmentation
- Vision-Language Cross-Attention for Real-Time Autonomous Driving
- Pulling Back the Curtain on Deep Networks
- CTG-Insight: A Multi-Agent Interpretable LLM Framework for Cardiotocography Analysis and Classification
- POLARIS: Explainable Artificial Intelligence for Mitigating Power Side-Channel Leakage
- On Explaining Visual Captioning with Hybrid Markov Logic Networks
- Compositional Function Networks: A High-Performance Alternative to Deep Neural Networks with Built-in Interpretability
- Self-learn to Explain Siamese Networks Robustly
- Towards Explainable Deep Clustering for Time Series Data
- Exposing the Illusion of Fairness: Auditing Vulnerabilities to Distributional Manipulation Attacks
- Towards trustworthy AI in materials mechanics through domain-guided attention
- Interpretable Anomaly-Based DDoS Detection in AI-RAN with XAI and LLMs
- Cross-Process Defect Attribution using Potential Loss Analysis
- Wafer Defect Root Cause Analysis with Partial Trajectory Regression
- Controllable Feature Whitening for Hyperparameter-Free Bias Mitigation
- Forest-Guided Clustering -- Shedding Light into the Random Forest Black Box
- Understanding Black-box Predictions via Influence Functions
- On the Performance of Concept Probing: The Influence of the Data (Extended Version)
- Concept Probing: Where to Find Human-Defined Concepts (Extended Version)
- CityHood: An Explainable Travel Recommender System for Cities and Neighborhoods
- Interpreting search result rankings through intent modeling
- Towards A Rigorous Science of Interpretable Machine Learning
- ProToken: Token-Level Attribution for Federated Large Language Models
- "Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
- Attri-Net: A Globally and Locally Inherently Interpretable Model for Multi-Label Classification Using Class-Specific Counterfactuals
- Poisoned classifiers are not only backdoored, they are fundamentally broken
Discussions
- “Why Should I Trust You?” Explaining the Predictions of Any Classifier [pdf] [hn, 3 points, 0 comments]
- [pdf] Explaining the Predictions of Any Classifier [hn, 1 points, 0 comments]
- Explaining black box models with LIME [hn, 1 points, 0 comments]
- “Why Should I Trust You?”: Explaining the Predictions of Any Classifier [hn, 1 points, 1 comments]
- Bringing trust into machine learning http://arxiv.org/abs/1602.04938 [bsky, 0 points, 0 comments]
Related