"Why Should I Trust You?": Explaining the Predictions of Any Classifier
2016/02/16 by Marco Túlio Ribeiro, Marco Tulio Ribeiro, Ribeiro, Marco Tulio +4 · 5 voices · 807 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Machine Learning and Data Classification #cs.AI #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1602.04938
arxiv published 2016/02/16 · arxiv created 2016/08/09 · arxiv updated 2016/08/10
Abstract
Despite widespread adoption, machine learning models remain mostly black boxes. Understanding the reasons behind predictions is, however, quite important in assessing trust, which is fundamental if one plans to take action based on a prediction, or when choosing whether to deploy a new model. Such understanding also provides insights into the model, which can be used to transform an untrustworthy model or prediction into a trustworthy one. In this work, we propose LIME, a novel explanation technique that explains the predictions of any classifier in an interpretable and faithful manner, by learning an interpretable model locally around the prediction. We also propose a method to explain models by presenting representative individual predictions and their explanations in a non-redundant way, framing the task as a submodular optimization problem. We demonstrate the flexibility of these methods by explaining different models for text (e.g. random forests) and image classification (e.g. neural networks). We show the utility of explanations via novel experiments, both simulated and with human subjects, on various scenarios that require trust: deciding if one should trust a prediction, choosing between models, improving an untrustworthy classifier, and identifying why a classifier should not be trusted.
Citations
Cited by
- Explainable Neural Inverse Kinematics for Obstacle-Aware Robotic Manipulation: A Comparative Analysis of IKNet Variants
- MiSS: A Logic-Driven Explanation of Minimal Sufficient Coalitions for Point Cloud Classifiers
- A CNN-Based Malaria Diagnosis from Blood Cell Images with SHAP and LIME Explainability
- Explainable Model Routing for Agentic Workflows
- LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
- Context-Aware Concept Distillation for Trustworthy Flood Prediction
- Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
- From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation
- Explainable Multimodal Regression via Information Decomposition
- Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations
- Beyond Adoption Intention How Trust in Augmented Analytics Relates to Perceived Decision Quality Among Non-Technical BI Users
- A Unified Framework for Uncertainty-Aware Explainable Artificial Intelligence: A Case Study in Power Quality Disturbance Classification
- Explainability Methods for Hardware Trojan Detection: A Systematic Comparison
- A Model of Causal Explanation on Neural Networks for Tabular Data
- Towards Generative Location Awareness for Disaster Response: A Probabilistic Cross-view Geolocalization Approach
- Explainable time-series forecasting with sampling-free SHAP for Transformers
- Toward Explaining Large Language Models in Software Engineering Tasks
- UbiQVision: Quantifying Uncertainty in XAI for Image Recognition
- Reason2Decide: Rationale-Driven Multi-Task Learning
- DecoKAN: Interpretable Decomposition for Forecasting Cryptocurrency Market Dynamics
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Augmenting Intelligence: A Hybrid Framework for Scalable and Stable Explanations
- Time-series Forecast for Indoor Zone Air Temperature with Long Horizons: A Case Study with Sensor-based Data from a Smart Building
- Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
- PROVEX: Enhancing SOC Analyst Trust with Explainable Provenance-Based IDS
- Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
- XAgen: An Explainability Tool for Identifying and Correcting Failures in Multi-Agent Workflows
- Interpretable Plant Leaf Disease Detection Using Attention-Enhanced CNN
- Adversarially Robust Detection of Harmful Online Content: A Computational Design Science Approach
- Explanation Beyond Intuition: A Testable Criterion for Inherent Explainability
- Enhancing Long Document Long Form Summarisation with Self-Planning
- UniCoMTE: A Universal Counterfactual Framework for Explaining Time-Series Classifiers on ECG Data
- Metanetworks as Regulatory Operators: Learning to Edit for Requirement Compliance
- Explainable AI in Big Data Fraud Detection
- Explaining the Reasoning of Large Language Models Using Attribution Graphs
- Multi-Agent Medical Decision Consensus Matrix System: An Intelligent Collaborative Framework for Oncology MDT Consultations
- Enhancing Interpretability for Vision Models via Shapley Value Optimization
- Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Stud
- Explainable reinforcement learning from human feedback to improve alignment
- Interpretable Hypothesis-Driven Trading:A Rigorous Walk-Forward Validation Framework for Market Microstructure Signals
- AgentSHAP: Interpreting LLM Agent Tool Importance with Monte Carlo Shapley Value Estimation
- Explainable Prediction of Economic Time Series Using IMFs and Neural Networks
- SRLR: Symbolic Regression based Logic Recovery to Counter Programmable Logic Controller Attacks
- Back to the Baseline: Examining Baseline Effects on Explainability Metrics
- Learning complete and explainable visual representations from itemized text supervision
- Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution
- Provable Recovery of Locally Important Signed Features and Interactions from Random Forest
- How Interpretable and Trustworthy are GAMs?
- Interpreto: An Explainability Library for Transformers
- Explainable Verification of Hierarchical Workflows Mined from Event Logs with Shapley Values
- TRUCE: TRUsted Compliance Enforcement Service for Secure Health Data Exchange
- SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning
- Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts
- Hybrid Attribution Priors for Explainable and Robust Model Training
- SSplain: Sparse and Smooth Explainer for Retinopathy of Prematurity Classification
- HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability
- ϕ-Table: A Statistical Explanation for Global SHAP
- A Geometric Unification of Concept Learning with Concept Cones
- CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
- Always Keep Your Promises: A Model-Agnostic Attribution Algorithm for Neural Networks
- CLUENet: Cluster Attention Makes Neural Networks Have Eyes
- Automated Plant Disease and Pest Detection System Using Hybrid Lightweight CNN-MobileViT Models for Diagnosis of Indigenous Crops
- MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution
- Improving Local Fidelity Through Sampling and Modeling Nonlinearity
- Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches
- Lyrics Matter: Exploiting the Power of Learnt Representations for Music Popularity Prediction
- Sell Data to AI Algorithms Without Revealing It: Secure Data Valuation and Sharing via Homomorphic Encryption
- Local surrogates for quantum machine learning
- Sample-Efficient Counterfactual Tuning for Compressor Pressure Control
- MASE: Interpretable NLP Models via Model-Agnostic Saliency Estimation
- Explainable AI for Smart Greenhouse Control: Interpretability of Temporal Fusion Transformer in the Internet of Robotic Things
- Out-of-the-box: Black-box Causal Attacks on Object Detectors
- A Framework for Causal Concept-based Model Explanations
- Beyond Additivity: Sparse Isotonic Shapley Regression toward Nonlinear Explainability
- Training Data Attribution for Image Generation using Ontology-Aligned Knowledge Graphs
- Demystifying Feature Engineering in Malware Analysis of API Call Sequences
- Forecasting India's Demographic Transition Under Fertility Policy Scenarios Using hybrid LSTM-PINN Model
- Supervised Contrastive Machine Unlearning of Background Bias in Sonar Image Classification with Fine-Grained Explainable AI
- Building Trustworthy AI for Materials Discovery: From Autonomous Laboratories to Z-scores
- Pushing the Boundaries of Interpretability: Incremental Enhancements to the Explainable Boosting Machine
- Neural Networks as Physics-Consistent Surrogates: An Explainable AI Validation Framework for Learning Constitutive Relations
- CodeFlowLM: Incremental Just-In-Time Defect Prediction with Pretrained Language Models and Exploratory Insights into Defect Localization
- From 'What-is' to 'What-if' in Human-Factor Analysis: A Post-Occupancy Evaluation Case
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model Predictions
- Faster Verified Explanations for Neural Networks
- DeFi TrustBoost: Blockchain and AI for Trustworthy Decentralized Financial Decisions
- Does the Model Say What the Data Says? A Simple Heuristic for Model Data Alignment
- Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring
- Analysis of Invasive Breast Cancer in Mammograms Using YOLO, Explainability, and Domain Adaptation
- BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning
- On Rank Graduation Metrics for High Dimensional Ordinal Data
- SHIC-XE: Viewpoint-Invariant Explainability via Dense 2D-3D Correspondences: an Application to Equine Pain Recognition
- Space Explanations of Neural Network Classification
- MATCH: Engineering Transparent and Controllable Conversational XAI Systems through Composable Building Blocks
- Physics-Informed Spiking Neural Networks via Conservative Flux Quantization
- The Directed Prediction Change - Efficient and Trustworthy Fidelity Assessment for Local Feature Attribution Methods
- Accumulated Local Effects and Graph Neural Networks for link prediction
- Ranking-Enhanced Anomaly Detection Using Active Learning-Assisted Attention Adversarial Dual AutoEncoders
- Generation, Evaluation, and Explanation of Novelists' Styles with Single-Token Prompts
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- Think Fast: Real-Time IoT Intrusion Reasoning Using IDS and LLMs at the Edge Gateway
- Matching-Based Few-Shot Semantic Segmentation Models Are Interpretable by Design
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
- Correlation-Aware Feature Attribution Based Explainable AI
- CID: Measuring Feature Importance Through Counterfactual Distributions
- MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm
- From Black Box to Insight: Explainable AI for Extreme Event Preparedness
- Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
- Multi-task GINN-LP for Multi-target Symbolic Regression
- From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability
- SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data
- LAYA: Layer-wise Attention Aggregation for Interpretable Depth-Aware Neural Networks
- Explainable Recommender Systems via Resolving Learning Representations
- Interpretable Fine-Gray Deep Survival Model for Competing Risks: Predicting Post-Discharge Foot Complications for Diabetic Patients in Ontario
- SCI: A Metacognitive Control for Signal Dynamics
- A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
- Additive Large Language Models for Semi-Structured Text
- ClinStructor: AI-Powered Structuring of Unstructured Clinical Texts
- Multi-omic Enriched Blood-Derived Digital Signatures Reveal Mechanistic and Confounding Disease Clusters for Differential Diagnosis
- Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
- Privacy-Preserving Explainable AIoT Application via SHAP Entropy Regularization
- When Thinking Pays Off: Incentive Alignment for Human-AI Collaboration
- MACIE: Multi-Agent Causal Intelligence Explainer for Collective Behavior Understanding
- llmSHAP: A Principled Approach to LLM Explainability
- Contrastive Integrated Gradients: A Feature Attribution-Based Method for Explaining Whole Slide Image Classification
- Improving Industrial Injection Molding Processes with Explainable AI for Quality Classification
- Increasing AI Explainability by LLM Driven Standard Processes
- Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
- Design Principles of Zero-Shot Self-Supervised Unknown Emitter Detectors
- Explainable Probabilistic Machine Learning for Predicting Drilling Fluid Loss of Circulation in Marun Oil Field
- Enhancing Adversarial Robustness of IoT Intrusion Detection via SHAP-Based Attribution Fingerprinting
- Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks
- DiagnoLLM: A Hybrid Bayesian Neural Language Framework for Interpretable Disease Diagnosis
- Adversarially Robust and Interpretable Magecart Malware Detection
- Unlocking the Black Box: A Five-Dimensional Framework for Evaluating Explainable AI in Credit Risk
- Uncovering Bugs in Formal Explainers: A Case Study with PyXAI
- Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
- Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
- Fair and Explainable Credit-Scoring under Concept Drift: Adaptive Explanation Frameworks for Evolving Populations
- Connecting the concepts of quantum state tomography and molecular representations for machine learning
- Geometry as a Missing Axis of Representation Quality: The Variational Geometric Information Bottleneck under Data Scarcity
- LLEXICORP: End-user Explainability of Convolutional Neural Networks
- AgentSLA : Towards a Service Level Agreement for AI Agents
- Community Detection on Model Explanation Graphs for Explainable AI
- Imbalanced Classification through the Lens of Spurious Correlations
- Multilingual BERT language model for medical tasks: Evaluation on domain-specific adaptation and cross-linguality
- Interpretable Model-Aware Counterfactual Explanations for Random Forest
- Before the Clinic: Transparent and Operable Design Principles for Healthcare AI
- Feature-Function Curvature Analysis: A Geometric Framework for Explaining Differentiable Models
- Counterfactual-based Agent Influence Ranker for Agentic AI Workflows
- Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
- Explainable AI: current status and future directions
- XRAI: Better Attributions Through Regions
- RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
- (EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations
- Comparing Rule-Based and Deep Learning Models for Patient Phenotyping
- On a Method to Measure Supervised Multiclass Model’s Interpretability: Application to Degradation Diagnosis (Short Paper)
- Reward Models are Metrics in a Trench Coat
- Convolutional neural networks in Vis–NIR chemometrics: From contradiction to conditional design
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- Deep representation learning for temporal inference in cancer omics: a systematic literature review
- Towards Robust Interpretability with Self-Explaining Neural Networks
- The many Shapley values for model explanation
- Neurodatascience: Past, Present, and Future
- Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
- Dating Apps and the Right to an Explanation
- Strategic inputs: feature selection from game-theoretic perspective
- FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning
- A Novel XAI-Enhanced Quantum Adversarial Networks for Velocity Dispersion Modeling in MaNGA Galaxies
- Feature-Guided Analysis of Neural Networks: A Replication Study
- AI in soil moisture remote sensing
- Toward Carbon-Neutral Human AI: Rethinking Data, Computation, and Learning Paradigms for Sustainable Intelligence
- MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
- Rule-Based Explanations for Retrieval-Augmented LLM Systems
- Towards Piece-by-Piece Explanations for Chess Positions with SHAP
- Tractable Shapley Values and Interactions via Tensor Networks
- From Black-box to Causal-box: Towards Building More Interpretable Models
- Additive Models Explained: A Computational Complexity Approach
- Embedding Explainable AI in NHS Clinical Safety: The Explainability-Enabled Clinical Safety Framework (ECSF)
- Knowledge-Driven Vision-Language Model for Plexus Detection in Hirschsprung's Disease
- Framework for Machine Evaluation of Reasoning Completeness in Large Language Models For Classification Tasks
- Citation Failure: Definition, Analysis and Efficient Mitigation
- Human-Centered LLM-Agent System for Detecting Anomalous Digital Asset Transactions
- ShapeX: Shapelet-Driven Post Hoc Explanations for Time Series Classification Models
- Green Finance and Carbon Emissions: A Nonlinear and Interaction Analysis Using Bayesian Additive Regression Trees
- From Prototypes to Sparse ECG Explanations: SHAP-Driven Counterfactuals for Multivariate Time-Series Multi-class Classification
- Inference on Variable Importance for Treatment Effect Heterogeneity: Shapley Values and Beyond
- Motion2Meaning: A Clinician-Centered Framework for Contestable LLM in Parkinson's Disease Gait Interpretation
- DSEBench: A Test Collection for Explainable Dataset Search with Examples
- Preliminary Quantitative Study on Explainability and Trust in AI Systems
- Reflections from Research Roundtables at the Conference on Health, Inference, and Learning (CHIL) 2025
- LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching
- Learnable Game-theoretic Policy Optimization for Data-centric Self-explanation Rationalization
- XD-RCDepth: Lightweight Radar-Camera Depth Estimation with Explainability-Aligned and Distribution-Aware Distillation
- Doctor Rashomon and the UNIVERSE of Madness: Variable Importance with Unobserved Confounding and the Rashomon Effect
- ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
- Unveiling the Vulnerability of Graph-LLMs: An Interpretable Multi-Dimensional Adversarial Attack on TAGs
- On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- Beyond single-model XAI: aggregating multi-model explanations for enhanced trustworthiness
- A Comprehensive Forecasting-Based Framework for Time Series Anomaly Detection: Benchmarking on the Numenta Anomaly Benchmark (NAB)
- Evaluating the Explainability of Vision Transformers in Medical Imaging
- Restricted Receptive Fields for Face Verification
- Assessing Policy Updates: Toward Trust-Preserving Intelligent User Interfaces
- f-INE: A Hypothesis Testing Framework for Estimating Influence under Training Randomness
- Denoised IPW-Lasso for Heterogeneous Treatment Effect Estimation in Randomized Experiments
- SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
- Explainable Human-in-the-Loop Segmentation via Critic Feedback Signals
- The Ethics Engine: A Modular Pipeline for Accessible Psychometric Assessment of Large Language Models
- Data Science in the Big Data Era: Analytics, Intelligence, and Future Challenges
- The Evolution of Artificial Intelligence Paradigms – Implications for Performance, Scalability, and Responsible Deployment
- Interpretable Generative and Discriminative Learning for Multimodal and Incomplete Clinical Data
- Towards Safer and Understandable Driver Intention Prediction
- It's 2025 -- Narrative Learning is the new baseline to beat for explainable machine learning
- From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening
- CARLE: A Hybrid Deep-Shallow Learning Framework for Robust and Explainable RUL Estimation of Rolling Element Bearings
- Training Feature Attribution for Vision Models
- Chain-of-Influence: Tracing Interdependencies Across Time and Features in Clinical Predictive Modelings
- Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
- Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
- Surrogate Graph Partitioning for Spatial Prediction
- IKNet: Interpretable Stock Price Prediction via Keyword-Guided Integration of News and Technical Indicators
- Machine Learning Algorithms for Financial Asset Price Forecasting
- Development of Mental Models in Human-AI Collaboration: A Conceptual Framework
- Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
- Enhancing Concept Localization in CLIP-based Concept Bottleneck Models
- Higher-Order Feature Attribution: Bridging Statistics, Explainable AI, and Topological Signal Processing
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
- The Role of Federated Learning in Improving Financial Security: A Survey
- QGraphLIME - Explaining Quantum Graph Neural Networks
- XAI-on-RAN: Explainable, AI-native, and GPU-Accelerated RAN Towards 6G
- Design Process of a Self Adaptive Smart Serious Games Ecosystem
- Reconsidering Requirements Engineering: Human-AI Collaboration in AI-Native Software Development
- Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
- Explainable but Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications
- Adaptive Node Feature Selection For Graph Neural Networks
- Onto-Epistemological Analysis of AI Explanations
- Provenance Networks: End-to-End Exemplar-Based Explainability
- Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
- Batch-CAM: Introduction to better reasoning in convolutional deep learning models
- Bayesian Neural Networks for Functional ANOVA model
- GDLNN: Marriage of Programming Language and Neural Networks for Accurate and Easy-to-Explain Graph Classification
- Explainable machine learning for multiscale thermal conductivity modeling in polymer nanocomposites with uncertainty quantification
- o-MEGA: Optimized Methods for Explanation Generation and Analysis
- Object-Centric Case-Based Reasoning via Argumentation
- MUSE-Explainer: Counterfactual Explanations for Symbolic Music Graph Classification Models
- ACT: Agentic Classification Tree
- ACE: Adapting sampling for Counterfactual Explanations
- Elucidating the Grey Atmosphere: SHAP Value Analysis of a Random Forest Atmospheric Neutral Density Model
- An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
- DeepProv: Behavioral Characterization and Repair of Neural Networks via Inference Provenance Graph Analysis
- Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- A Fuzzy Logic-Based Framework for Explainable Machine Learning in Big Data Analytics
- Identifying Information-Transfer Nodes in a Recurrent Neural Network Reveals Dynamic Representations
- Understanding Collaboration between Professional Designers and Decision-making AI: A Case Study in the Workplace
- Fidelity-Aware Data Composition for Robust Robot Generalization
- Fast and accurate explanations of distance-based classifiers by uncovering latent explanatory structures
- Improving Simple Models with Confidence Profiles
- Deep learning-based morphological classification of ceramics: A case study of 3D point cloud analysis for Sue ware, Japan
- Who Trusts AI for Health Information? A Cross-National Survey on Trust Determinants in Four European Countries
- On The Variability of Concept Activation Vectors
- A Novel Hybrid Deep Learning and Chaotic Dynamics Approach for Thyroid Cancer Classification
- CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
- SHAPoint: Task-Agnostic, Efficient, and Interpretable Point-Based Risk Scoring via Shapley Values
- Accuracy-Robustness Trade Off via Spiking Neural Network Gradient Sparsity Trail
- ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series Forecasting
- Towards Human-interpretable Explanation in Code Clone Detection using LLM-based Post Hoc Explainer
- MindCraft: How Concept Trees Take Shape In Deep Models
- Learning Temporal Saliency for Time Series Forecasting with Cross-Scale Attention
- CausalKANs: interpretable treatment effect estimation with Kolmogorov-Arnold networks
- Explaining multimodal LLMs via intra-modal token interactions
- Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
- Bridging the Gap Between Scientific Laws Derived by AI Systems and Canonical Knowledge via Abductive Inference with AI-Noether
- Optimal Robust Recourse with Lp-Bounded Model Change
- Explaining Fine Tuned LLMs via Counterfactuals A Knowledge Graph Driven Framework
- Learning Conformal Explainers for Image Classifiers
- EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email Defense
- GenFacts-Generative Counterfactual Explanations for Multi-Variate Time Series
- ExpIDS: A Drift-adaptable Network Intrusion Detection System With Improved Explainability
- Investigating Modality Contribution in Audio LLMs for Music
- Exact Subgraph Isomorphism Network with Mixed L0,2 Norm Constraint for Predictive Graph Mining
- TSKAN: Interpretable Machine Learning for QoE modeling over Time Series Data
- Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning
- Every Character Counts: From Vulnerability to Defense in Phishing Detection
- Trust and Transparency in AI: Industry Voices on Data, Ethics, and Compliance
- HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data
- Improving Mental Health Screening and Early Risk Detection in Spanish
- Class-Aware Reinforcement Learning for Counterfactual Explanation Generation
- FADEx: Feature Attribution and Distortion-based Explanation of Dimensionality Reduction
- Quantitative Evaluations on Saliency Methods: An Experimental Study
- Explainable AI Isn't Enough! Rethinking Algorithmic Contestability
- Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments
- Physics-constrained neural ordinary differential equation models to discover and predict microbial community dynamics
- Concerning Uncertainty -- A Systematic Survey of Uncertainty-Aware XAI
- Towards Transparent Application of Machine Learning in Video Processing
- FA(IR)2MA-GLVQ – A hidden-feature-bias mitigation approach for fairness in classification learning based on generalized matrix learning vector quantization
- Toward Human-Centered Explainability: Natural Language Explanations for Anomaly Detection
- Artificial Intelligence-Powered Raman Spectroscopy through Open Science and FAIR Principles
- Explaining and interpreting hyperdimensional computing classifiers on tabular data
- Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
- Learning From Simulators: A Theory of Simulation-Grounded Learning
- Path-Weighted Integrated Gradients for Interpretable Dementia Classification
- Personalized Prediction By Learning Halfspace Reference Classes Under Well-Behaved Distribution
- Building Transparency in Deep Learning-Powered Network Traffic Classification: A Traffic-Explainer Framework
- Attention Consistency for LLMs Explanation
- Looking in the mirror: A faithful counterfactual explanation method for interpreting deep image classification models
- Towards a Transparent and Interpretable AI Model for Medical Image Classifications
- Interpretable Clinical Classification with Kolmogorov-Arnold Networks
- CausalSent: Interpretable Sentiment Classification with RieszNet
- Benchmarking Class Activation Map Methods for Explainable Brain Hemorrhage Classification on Hemorica Dataset
- Learning Interpretable Differentiable Logic Networks for Time-Series Classification
- RulER: Automated Rule-Based Semantic Error Localization and Repair for Code Translation
- Copycat vs. Original: Multi-modal Pretraining and Variable Importance in Box-office Prediction
- NeRF-based Visualization of 3D Cues Supporting Data-Driven Spacecraft Pose Estimation
- Attention Lattice Adapter: Visual Explanation Generation for Visual Foundation Model
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- LLM on a Budget: Active Knowledge Distillation for Efficient Classification of Large Text Corpora
- Out of Distribution Detection in Self-adaptive Robots with AI-powered Digital Twins
- L-XAIDS: A LIME-based eXplainable AI framework for Intrusion Detection Systems
- Explainable Counterfactual Reasoning in Depression Medication Selection at Multi-Levels (Personalized and Population)
- An empirical comparison of machine learning methods for text-based sentiment analysis of online consumer reviews
- Regulating genome language models: navigating policy challenges at the intersection of AI and genetics
- Evalet: Evaluating Large Language Models through Functional Fragmentation
- Clarifying Model Transparency: Interpretability versus Explainability in Deep Learning with MNIST and IMDB Examples
- CUBE: Contrastive Understanding by Balanced Experiments
- Progressive Disclosure: Designing for Effective Transparency
- NuGraph2 with Explainability: Post-hoc Explanations for Geometric Neural Network Predictions
- Can Explainable AI Explain Unfairness? A Framework for Evaluating Explainable AI
- Model-agnostic post-hoc explainability for recommender systems
- Investigating Feature Attribution for 5G Network Intrusion Detection
- Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models
- Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
- Evaluation of Black-Box XAI Approaches for Predictors of Values of Boolean Formulae
- An Autoencoder and Vision Transformer-based Interpretability Analysis of the Differences in Automated Staging of Second and Third Molars
- Database Views as Explanations for Relational Deep Learning
- STRIDE: Subset-Free Functional Decomposition for XAI in Tabular Settings
- Uncertainty Awareness and Trust in Explainable AI- On Trust Calibration using Local and Global Explanations
- Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"
- An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
- Explainability of CNN Based Classification Models for Acoustic Signal
- Minimal Data, Maximum Clarity: A Heuristic for Explaining Optimization
- An Interpretable Deep Learning Model for General Insurance Pricing
- Transparency of medical artificial intelligence systems
- Hybrid GCN-GRU Model for Anomaly Detection in Cryptocurrency Transactions
- MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
- Expert-Guided Explainable Few-Shot Learning for Medical Image Diagnosis
- Rule-Based Moral Principles for Explaining Uncertainty in Natural Language Generation
- An Explainable Deep Neural Network with Frequency-Aware Channel and Spatial Refinement for Flood Prediction in Sustainable Cities
- A Survey of Challenges and Opportunities in Sensing and Analytics for Cardiovascular Disorders
- An Approach to Grounding AI Model Evaluations in Human-derived Criteria
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- BIDO: An Out-Of-Distribution Resistant Image-based Malware Detector
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- Systematic Evaluation of Attribution Methods: Eliminating Threshold Bias and Revealing Method-Dependent Performance Patterns
- Enhancing Interpretability and Effectiveness in Recommendation with Numerical Features via Learning to Contrast the Counterfactual samples
- A XAI-based Framework for Frequency Subband Characterization of Cough Spectrograms in Chronic Respiratory Disease
- Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring
- Who Owns The Robot?: Four Ethical and Socio-technical Questions about Wellbeing Robots in the Real World through Community Engagement
- An Information-Flow Perspective on Explainability Requirements: Specification and Verification
- Wrong Model, Right Uncertainty: Spatial Associations for Discrete Data with Misspecification
- DRetNet: A Novel Deep Learning Framework for Diabetic Retinopathy Diagnosis
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- An Explainable Gaussian Process Auto-encoder for Tabular Data
- Causal SHAP: Feature Attribution with Dependency Awareness through Causal Discovery
- Tabular Diffusion Counterfactual Explanations
- XAI-Driven Machine Learning System for Driving Style Recognition and Personalized Recommendations
- Large Language Model Integration with Reinforcement Learning to Augment Decision-Making in Autonomous Cyber Operations
- GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability
- Individualized and Interpretable Sleep Forecasting via a Two-Stage Adaptive Spatial-Temporal Model
- Axiomatic Attribution for Deep Networks
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating
- Multi-View Graph Convolution Network for Internal Talent Recommendation Based on Enterprise Emails
- Automated Quality Assessment for LLM-Based Complex Qualitative Coding: A Confidence-Diversity Framework
- Safer Skin Lesion Classification with Global Class Activation Probability Map Evaluation and SafeML
- A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
- Attention-Based Explainability for Structure-Property Relationships
- Breaking the Black Box: Inherently Interpretable Physics-Constrained Machine Learning With Weighted Mixed-Effects for Imbalanced Seismic Data
- VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- PAX-TS: Model-agnostic multi-granular explanations for time series forecasting via localized perturbations
- How Reliable are LLMs for Reasoning on the Re-ranking task?
- Locally Pareto-Optimal Interpretations for Black-Box Machine Learning Models
- The Loupe: A Plug-and-Play Attention Module for Amplifying Discriminative Features in Vision Transformers
- QA-VLM: Providing human-interpretable quality assessment for wire-feed laser additive manufacturing parts with Vision Language Models
- XAI-Driven Spectral Analysis of Cough Sounds for Respiratory Disease Characterization
- Exact Shapley Attributions in Quadratic-time for FANOVA Gaussian Processes
- A Comparative Evaluation of Teacher-Guided Reinforcement Learning Techniques for Autonomous Cyber Operations
- Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
- ITL-LIME: Instance-Based Transfer Learning for Enhancing Local Explanations in Low-Resource Data Settings
- Compressed Models are NOT Trust-equivalent to Their Large Counterparts
- Interpreting Time Series Forecasts with LIME and SHAP: A Case Study on the Air Passengers Dataset
- Design and Validation of a Responsible Artificial Intelligence-based System for the Referral of Diabetic Retinopathy Patients
- Rigorous Feature Importance Scores based on Shapley Value and Banzhaf Index
- Informative Post-Hoc Explanations Only Exist for Simple Functions
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- When Explainability Meets Privacy: An Investigation at the Intersection of Post-hoc Explainability and Differential Privacy in the Context of Natural Language Processing
- RealAC: A Domain-Agnostic Framework for Realistic and Actionable Counterfactual Explanations
- An unexpected unity among methods for interpreting model predictions
- Explainable AI Technique in Lung Cancer Detection Using Convolutional Neural Networks
- Extending the Entropic Potential of Events for Uncertainty Quantification and Decision-Making in Artificial Intelligence
- Adoption of Explainable Natural Language Processing: Perspectives from Industry and Academia on Practices and Challenges
- A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- Understanding Dementia Speech Alignment with Diffusion-Based Image Generation
- SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling
- Reveal-Bangla: A Dataset for Cross-Lingual Multi-Step Reasoning Evaluation
- Beyond Technocratic XAI: The Who, What & How in Explanation Design
- Hierarchical Variable Importance with Statistical Control for Medical Data-Based Prediction
- Beyond Predictions: A Study of AI Strength and Weakness Transparency Communication on Human-AI Collaboration
- Deep Reinforcement Learning with Local Interpretability for Transparent Microgrid Resilience Energy Management
- Exploring Content and Social Connections of Fake News with Explainable Text and Graph Learning
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
- Feature contributions and predictive modeling of aeolian sand transport detection in the atmospheric surface layer
- An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal
- AttnTrace: Attention-based Context Traceback for Long-Context LLMs
- SHAP Stability in Credit Risk Management: A Case Study in Credit Card Default Model
- TeSent: A Benchmark Dataset for Fairness-aware Explainable Sentiment Classification in Telugu
- AutoSIGHT: Automatic Eye Tracking-based System for Immediate Grading of Human experTise
- TofuML: A Spatio-Physical Interactive Machine Learning Device for Interactive Exploration of Machine Learning for Novices
- MetaExplainer: A Framework to Generate Multi-Type User-Centered Explanations for AI Systems
- Sufficient, Necessary and Complete Causal Explanations in Image Classification
- Your Model Is Unfair, Are You Even Aware? Inverse Relationship Between Comprehension and Trust in Explainability Visualizations of Biased ML Models
- TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning
- Teaching the Teacher: Improving Neural Network Distillability for Symbolic Regression via Jacobian Regularization
- XAI for Point Cloud Data using Perturbations based on Meaningful Segmentation
- Vision-Language Cross-Attention for Real-Time Autonomous Driving
- Pulling Back the Curtain on Deep Networks
- CTG-Insight: A Multi-Agent Interpretable LLM Framework for Cardiotocography Analysis and Classification
- POLARIS: Explainable Artificial Intelligence for Mitigating Power Side-Channel Leakage
- On Explaining Visual Captioning with Hybrid Markov Logic Networks
- Compositional Function Networks: A High-Performance Alternative to Deep Neural Networks with Built-in Interpretability
- Self-learn to Explain Siamese Networks Robustly
- Towards Explainable Deep Clustering for Time Series Data
- Exposing the Illusion of Fairness: Auditing Vulnerabilities to Distributional Manipulation Attacks
- Towards trustworthy AI in materials mechanics through domain-guided attention
- Interpretable Anomaly-Based DDoS Detection in AI-RAN with XAI and LLMs
- Cross-Process Defect Attribution using Potential Loss Analysis
- Wafer Defect Root Cause Analysis with Partial Trajectory Regression
- Controllable Feature Whitening for Hyperparameter-Free Bias Mitigation
- Explainability in machine learning: a pedagogical perspective
- Advancing Responsible Innovation in Agentic AI: A study of Ethical Frameworks for Household Automation
- Forest-Guided Clustering -- Shedding Light into the Random Forest Black Box
- Understanding Black-box Predictions via Influence Functions
- On the Performance of Concept Probing: The Influence of the Data (Extended Version)
- Concept Probing: Where to Find Human-Defined Concepts (Extended Version)
- CityHood: An Explainable Travel Recommender System for Cities and Neighborhoods
- Interpreting search result rankings through intent modeling
- XplainAct: Visualization for Personalized Intervention Insights
- Fraud is Not Just Rarity: A Causal Prototype Attention Approach to Realistic Synthetic Oversampling
- Towards A Rigorous Science of Interpretable Machine Learning
- ProToken: Token-Level Attribution for Federated Large Language Models
- Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
- XAI-Guided Analysis of Residual Networks for Interpretable Pneumonia Detection in Paediatric Chest X-rays
- "Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
- Confidence-Filtered Relevance (CFR): An Interpretable and Uncertainty-Aware Machine Learning Framework for Naturalness Assessment in Satellite Imagery
- InTraVisTo: Inside Transformer Visualisation Tool
- Attri-Net: A Globally and Locally Inherently Interpretable Model for Multi-Label Classification Using Class-Specific Counterfactuals
- Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics
- Poisoned classifiers are not only backdoored, they are fundamentally broken
- A Computational Framework to Identify Self-Aspects in Text
- Knowing What You Cannot Explain: Learning to Reject Low-Quality Explanations
- Robust Explanations Through Uncertainty Decomposition: A Path to Trustworthier AI
- MUPAX: Multidimensional Problem Agnostic eXplainable AI
- Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation
- Mapping Emotions in the Brain: A Bi-Hemispheric Neural Model with Explainable Deep Learning
- Explainable Evidential Clustering
- OrdShap: Feature Position Importance for Sequential Black-Box Models
- Explainable AI for online disinformation detection: Insights from a design science research project
- AURA: A Multi-Agent Intelligence Framework for Knowledge-Enhanced Cyber Threat Attribution
- ClarifAI: Enhancing AI Interpretability and Transparency through Case-Based Reasoning and Ontology-Driven Approach for Improved Decision-Making
- Explainable Compliance Detection with Multi-Hop Natural Language Inference on Assurance Case Structure
- Explaining, Fast and Slow: Abstraction and Refinement of Provable Explanations
- Contestability in Quantitative Argumentation
- Impact of Accuracy on Model Interpretations
- Unveiling trust in AI: the interplay of antecedents, consequences, and cultural dynamics
- TaylorPODA: A Taylor Expansion-Based Method to Improve Post-Hoc Attributions for Opaque Models
- Towards Applying Large Language Models to Complement Single-Cell Foundation Models
- The Authority of "Fair" in Machine Learning
- Winsor-CAM: Human-Tunable Visual Explanations from Deep Networks via Layer-Wise Winsorization
- Conditional Chemical Language Models are Versatile Tools in Drug Discovery
- Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications
- Towards Interpretable Multi-Task Learning Using Bilevel Programming
- PLEX: Perturbation-free Local Explanations for LLM-Based Text Classification
- From Classical Machine Learning to Emerging Foundation Models: Review on Multimodal Data Integration for Cancer Research
- NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
- Bridging the gap in FER: addressing age bias in deep learning
- On Trustworthy Rule-Based Models and Explanations
- Bluish Veil Detection and Lesion Classification using Custom Deep Learnable Layers with Explainable Artificial Intelligence (XAI)
- Comprehensive Evaluation of Prototype Neural Networks
- Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey
- Understanding Malware Propagation Dynamics through Scientific Machine Learning
- Explaining Surface Layer Theory Departures in Marine Flux Profiles with Data-Driven Discovery
- Feature-Guided Neighbor Selection for Non-Expert Evaluation of Model Predictions
- Enhancing the Interpretability of Rule-based Explanations through Information Retrieval
- Taming Data Challenges in ML-based Security Tasks Using Generative AI
- On the Effectiveness of Methods and Metrics for Explainable AI in Remote Sensing Image Scene Classification
- Machine Learning from Explanations
- What Is Missing in Surgical Risk Stratification and Outcome Prediction: A Scoping Review of End-to-End Machine Learning Approaches
- Towards integration of Privacy Enhancing Technologies in Explainable Artificial Intelligence
- Machine Learning in Acoustics: A Review and Open-Source Repository
- Attributing Data for Sharpness-Aware Minimization
- RareSense: Rarity-Aware Similarity Search for Anomaly Retrieval in Transactional Data
- Unraveling the Black Box of Neural Networks: A Dynamic Extremum Mapper
- Interpretable Diffusion Models with B-cos Networks
- Explainable Information Retrieval in the Audit Domain
- DeltaSHAP: Explaining Prediction Evolutions in Online Patient Monitoring with Shapley Values
- Deep learning-based stacked models for cyber-attack detection in industrial internet of things
- Benchmarking the Discovery Engine
- Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
- Towards a Signal Detection Based Measure for Assessing Information Quality of Explainable Recommender Systems
- Two-Stage Reasoning-Infused Learning: Improving Classification with LLM-Generated Reasoning
- Emerging AI Approaches for Cancer Spatial Omics
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
- We Are AI: Taking Control of Technology
- A versatile XAI-based framework for efficient and explainable intrusion detection systems
- AICO: Feature Significance Tests for Supervised Learning
- Interpretable by Design: MH-AutoML for Transparent and Efficient Android Malware Detection without Compromising Performance
- Token Activation Map to Visually Explain Multimodal LLMs
- Aggregating Local Saliency Maps for Semi-Global Explainable Image Classification
- Evaluating explainable AI for deep learning-based network intrusion detection system alert classification
- Evaluating Pavement Deterioration Rates Due to Flooding Events Using Explainable AI
- Predicting and Explaining Customer Data Sharing in the Open Banking
- GRASP-PsONet: Gradient-based Removal of Spurious Patterns for PsOriasis Severity Classification
- Towards Large Language Models with Self-Consistent Natural Language Explanations
- IXAII: An Interactive Explainable Artificial Intelligence Interface for Decision Support Systems
- Towards Transparent AI: A Survey on Explainable Large Language Models
- Understanding plant phenotypes in crop breeding through explainable AI
- Explainable AI for Radar Resource Management: Modified LIME in Deep Reinforcement Learning
- Communicating Smartly in Molecular Communication Environments: Neural Networks in the Internet of Bio-Nano Things
- Domain Knowledge in Artificial Intelligence: Using Conceptual Modeling to Increase Machine Learning Accuracy and Explainability
- Autonomous Cyber Resilience via a Co-Evolutionary Arms Race within a Fortified Digital Twin Sandbox
- SEZ-HARN: Self-Explainable Zero-shot Human Activity Recognition Network
- Context Attribution with Multi-Armed Bandit Optimization
- How trust networks shape students' opinions about the proficiency of artificially intelligent assistants
- Toward the Explainability of Protein Language Models
- Explaining deep neural network models for electricity price forecasting with XAI
- Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
- Why Uncertainty Calibration Matters for Reliable Perturbation-based Explanations
- Model Guidance via Robust Feature Attribution
- FREQuency ATTribution: Benchmarking Frequency-based Occlusion for Time Series Data
- Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications
- Pathwise Explanation of ReLU Neural Networks
- Decoding Federated Learning: The FedNAM+ Conformal Revolution
- Generative Grasp Detection and Estimation with Concept Learning-based Safety Criteria
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
- AQUA20: A Benchmark Dataset for Underwater Species Classification under Challenging Conditions
- Anomaly Detection in Event-triggered Traffic Time Series via Similarity Learning
- Uncertainty-Aware Information Pursuit for Interpretable and Reliable Medical Image Analysis
- A novel fusion architecture for detecting Parkinson’s Disease using semi-supervised speech embeddings
- From Lab to Factory: Pitfalls and Guidelines for Self-/Unsupervised Defect Detection on Low-Quality Industrial Images
- The Role of Explanation Styles and Perceived Accuracy on Decision Making in Predictive Process Monitoring
- A Hybrid DeBERTa and Gated Broad Learning System for Cyberbullying Detection in English Text
- TRUST: Transparent, Robust and Ultra-Sparse Trees
- Statistical Hypothesis Testing for Auditing Robustness in Language Models
- Pixel-level Certified Explanations via Randomized Smoothing
- A Federated Random Forest Solution for Secure Distributed Machine Learning
- Deep Learning Approach for Generating MRA Images From 3D Quantitative Synthetic MRI Without Additional Scans
- Leveraging Predictive Equivalence in Decision Trees
- Model compression using knowledge distillation with integrated gradients
- Mxplainer: Explain and Learn Insights by Imitating Mahjong Agents
- A Comprehensive Analysis of COVID-19 Detection Using Bangladeshi Data and Explainable AI
- AutoSAS: a new human-aside-the-loop paradigm for automated SAS fitting for high throughput and autonomous experimentation
- Rethinking Explainability in the Era of Multimodal AI
- SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
- VARSHAP: Addressing Global Dependency Problems in Explainable AI with Variance-Based Local Feature Attribution
- "Faithful to What?" On the Limits of Fidelity-Based Explanations
- Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI
- Because we have LLMs, we Can and Should Pursue Agentic Interpretability
- AI-Based Software Vulnerability Detection: A Systematic Literature Review
- Reasoning Models Don't Always Say What They Think
- WISCA: A Consensus-Based Approach to Harmonizing Interpretability in Tabular Datasets
- Explainability in Context: A Multilevel Framework Aligning AI Explanations with Stakeholder with LLMs
- Machine learning and behavioral economics for personalized choice architecture
- There's Waldo: PCB Tamper Forensic Analysis using Explainable AI on Impedance Signatures
- Personalized Interpretability -- Interactive Alignment of Prototypical Parts Networks
- Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
- A Unified Framework for Provably Efficient Algorithms to Estimate Shapley Values
- Modern approaches to building interpretable models of the property market using machine learning on the base of mass cadastral valuation
- Deciding When Not to Decide: Indeterminacy-Aware Intrusion Detection with NeutroSENSE
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
- TIMING: Temporality-Aware Integrated Gradients for Time Series Explanation
- NIMO: a Nonlinear Interpretable MOdel
- Recent Advances in Medical Image Classification
- A Comprehensive Study on Medical Image Segmentation using Deep Neural Networks
- DURA-CPS: A Multi-Role Orchestrator for Dependability Assurance in LLM-Enabled Cyber-Physical Systems
- Benchmarking Time-localized Explanations for Audio Classification Models
- TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- Exploring Convolutional Neural Networks for Rice Grain Classification: An Explainable AI Approach
- Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable AI Auditors
- Prediction via Shapley Value Regression
- The Promise and Peril of Human Evaluation for Model Interpretability
- Deep Learning for Computational Chemistry
- Explainable AI: XAI-Guided Context-Aware Data Augmentation
- Causal Explanations Over Time: Articulated Reasoning for Interactive Environments
- Explainability-Based Token Replacement on LLM-Generated Text
- A Unifying Bias-aware Multidisciplinary Framework for Investigating Socio-Technical Issues
- Identifying Alzheimer's Disease Prediction Strategies of Convolutional Neural Network Classifiers using R2* Maps and Spectral Clustering
- TracLLM: A Generic Framework for Attributing Long Context LLMs
- Interpretabilité des modèles : état des lieux des méthodes et application à l'assurance
- SynLang and Symbiotic Epistemology: A Manifesto for Conscious Human-AI Collaboration
- On the Necessity of Multi-Domain Explanation: An Uncertainty Principle Approach for Deep Time Series Models
- Composable Building Blocks for Controllable and Transparent Interactive AI Systems
- Towards Better Generalization and Interpretability in Unsupervised Concept-Based Models
- XAI-Units: Benchmarking Explainability Methods with Unit Tests
- ExplainBench: A Benchmark Framework for Local Model Explanations in Fairness-Critical Applications
- C2G-Net: Exploiting Morphological Properties for Image Classification
- Deep Learning in Information Security
- Feature Attribution from First Principles
- Interpretable phenotyping of Heart Failure patients with Dutch discharge letters
- Feature Relevancy, Necessity and Usefulness: Complexity and Algorithms
- Multi-criteria Rank-based Aggregation for Explainable AI
- In generative artificial intelligence we trust: unpacking determinants and outcomes for cognitive trust
- Learning Interpretable Differentiable Logic Networks for Tabular Regression
- Enhancing Uncertainty Estimation and Interpretability via Bayesian Non-negative Decision Layer
- From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications
- Effective Context in Neural Speech Models
- A New Approach to Backtracking Counterfactual Explanations: A Unified Causal Framework for Efficient Model Interpretability
- Efficient Preimage Approximation for Neural Network Certification
- Counterfactual Multi-player Bandits for Explainable Recommendation Diversification
- Privacy Risks and Preservation Methods in Explainable Artificial Intelligence: A Scoping Review
- Humble AI in the real-world: the case of algorithmic hiring
- MultiPhishGuard: An LLM-based Multi-Agent System for Phishing Email Detection
- Comparing Neural Network Encodings for Logic-based Explainability
- Explanations of Machine Learning predictions: a mandatory step for its application to Operational Processes
- Class Introspection: A Novel Technique for Detecting Unlabeled Subclasses by Leveraging Classifier Explainability Methods
- InfoCons: Identifying Interpretable Critical Concepts in Point Clouds via Information Theory
- Principles of Explanation in Human-AI Systems
- Robustness questions the interpretability of graph neural networks: what to do?
- A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
- AI for Regulatory Affairs: Balancing Accuracy, Interpretability, and Computational Cost in Medical Device Classification
- CRITS: Convolutional Rectifier for Interpretable Time Series Classification
- Towards Uncertainty Aware Task Delegation and Human-AI Collaborative Decision-Making
- Emerging categories in scientific explanations
- ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
- EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion
- Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability
- Reliable Deep Grade Prediction with Uncertainty Estimation
- On the reliability of feature attribution methods for speech classification
- The Bidirectional Relationship Between XAI and Regulation: Operationalizing XAI for the AI Act
- Comprehensive Lung Disease Detection Using Deep Learning Models and Hybrid Chest X-ray Data with Explainable AI
- Representation Learning for Electronic Health Records
- Explainable embeddings with Distance Explainer
- Does Explainability Transfer? A Controlled Benchmark of Attribution Methods on Vision Transformers and CNNs
- Explaining Unreliable Perception in Automated Driving: A Fuzzy-based Monitoring Approach
- The Evolution of Alpha in Finance Harnessing Human Insight and LLM Agents
- BACON: A fully explainable AI model with graded logic for decision making problems
- Explainable AI for Securing Healthcare in IoT-Integrated 6G Wireless Networks
- CSAGC-IDS: A Dual-Module Deep Learning Network Intrusion Detection Model for Complex and Imbalanced Data
- Enhancing Interpretability of Sparse Latent Representations with Class Information
- From Representation to Mediation: A New Agenda for Conceptual Modeling Research in A Digital World
- EPIC: Explanation of Pretrained Image Classification Networks via Prototype
- Dementia Through Different Eyes: Explainable Modeling of Human and LLM Perceptions for Early Awareness
- VGER: Voxel-Guided Global Event Ranking for Event Cloud Attribution
- Towards Budget-Friendly Model-Agnostic Explanation Generation for Large Language Models
- Towards Structurally Explainable Machine-Generated Text Detection: A Graph-Perspective Framework
- Fixed Point Explainability
- Attribution Projection Calculus: A Novel Framework for Causal Inference in Bayesian Networks
- LAMP: Extracting Locally Linear Decision Surfaces from LLM World Models
- AdaptMol: Adaptive Fusion from Sequence String to Topological Structure for Few-shot Drug Discovery
- Heart2Mind: Human-Centered Contestable Psychiatric Disorder Diagnosis System using Wearable ECG Monitors
- Analysis of Customer Journeys Using Prototype Detection and Counterfactual Explanations for Sequential Data
- Most General Explanations of Tree Ensembles (Extended Version)
- Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP
- Concept-Guided Interpretability via Neural Chunking
- Explaining Strategic Decisions in Multi-Agent Reinforcement Learning for Aerial Combat Tactics
- CUBIC: Concept Embeddings for Unsupervised Bias Identification using VLMs
- WebXAII: an open-source web framework to study human-XAI interaction
- PnPXAI: A Universal XAI Framework Providing Automatic Explanations Across Diverse Modalities and Models
- Explainable Hybrid Feature Selection for Intrusion Detection in Internet of Medical Things Environments
- A User Study Evaluating Argumentative Explanations in Diagnostic Decision Support
- Financial Fraud Detection Using Explainable AI and Stacking Ensemble Methods
- The Impact of Climatic Factors on Respiratory Pharmaceutical Demand: A Comparison of Forecasting Models for Greece
- Explainability Through Human-Centric Design for XAI in Lung Cancer Detection
- Cybersecurity threat detection based on a UEBA framework using Deep Autoencoders
- Rhetorical XAI: Explaining AI's Benefits as well as its Use via Rhetorical Design
- Implet: A Post-hoc Subsequence Explainer for Time Series Models
- Integrating Natural Language Processing and Exercise Monitoring for Early Diagnosis of Metabolic Syndrome: A Deep Learning Approach
- A Deep Learning-Driven Inhalation Injury Grading Assistant Using Bronchoscopy Images
- AI and Generative AI Transforming Disaster Management: A Survey of Damage Assessment and Response Techniques
- From Search To Sampling: Generative Models For Robust Algorithmic Recourse
- Concept-Level Explainability for Auditing & Steering LLM Responses
- Interpretable Event Diagnosis in Water Distribution Networks
- DocVXQA: Context-Aware Visual Explanations for Document Question Answering
- Discovering Concept Directions from Diffusion-based Counterfactuals via Latent Clustering
- A Survey on Foundation Models for Personalized Federated Intelligence
- Navigating the Rashomon Effect: How Personalization Can Help Adjust Interpretable Machine Learning Models to Individual Users
- Explainable AI the Latest Advancements and New Trends
- Integrating Explainable AI in Medical Devices: Technical, Clinical and Regulatory Insights and Recommendations
- See What I Mean? CUE: A Cognitive Model of Understanding Explanations
- Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation
- What Do People Want to Know About Artificial Intelligence (AI)? The Importance of Answering End-User Questions to Explain Autonomous Vehicle (AV) Decisions
- From Pixels to Perception: Interpretable Predictions via Instance-wise Grouped Feature Selection
- Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
- Human in the Latent Loop (HILL): Interactively Guiding Model Training Through Human Intuition
- PointExplainer: Towards Transparent Parkinson's Disease Diagnosis
- Explainable Face Recognition via Improved Localization
- Interpretable graph-based models on multimodal biomedical data integration: A technical review and benchmarking
- Understanding the Mechanisms Behind Structural Influences on Link Prediction: A Case Study on FB15k-237
- ABE: A Unified Framework for Robust and Faithful Attribution-Based Explainability
- EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
- Explainable Machine Learning for Cyberattack Identification from Traffic Flows
- Thinking Outside the Template with Modular GP-GOMEA
- Exploring the Impact of Explainable AI and Cognitive Capabilities on Users' Decisions
- Machine Learning for Cyber-Attack Identification from Traffic Flows
- Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
- Overview and practical recommendations on using Shapley Values for identifying predictive biomarkers via CATE modeling
- Combining LLMs with Logic-Based Framework to Explain MCTS
- A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
- Measuring Explainer Stability via Attribution Separability
- Interpretation of Prediction Models Using the Input Gradient
- Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
- Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies
- Validating Causal Abstraction Metrics on Simulated Complex Systems
- Self-Ablating Transformers: More Interpretability, Less Sparsity
- Computational Identification of Regulatory Statements in EU Legislation
- Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery
- Interpreting "Interpretability" and Explaining "Explainability" in Machine Learning in Physics
- Embodied Explainability and Ontological Obstacles: Why We Struggle to Explain the Answers of Large Language Models (LLMs)
- Martingale Doppelgänger-Eval: An Identification Framework for Auditing Candlestick Understanding in Vision-Language Models
- Explainable Systematic Analysis for Synthetic Aperture Sonar Imagery
- The Perceptual Gap: Why We Need Accessible XAI for Assistive Technologies
- Explainability and justification of automatic-decision making: A conceptual framework and a practical application
- IP-CRR: Information Pursuit for Interpretable Classification of Chest Radiology Reports
- ArrhythmiaVision: Resource-Conscious Deep Learning Models with Visual Explanations for ECG Arrhythmia Classification
- Could Large Language Models work as Post-hoc Explainability Tools in Credit Risk Models?
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- Measuring Understanding Through Discrete Compositional Knowledge Structures in Hierarchical Automata
- In defence of post-hoc explanations in medical AI
- Unsupervised Surrogate Anomaly Detection
- RuleKit 2: Faster and simpler rule learning
- Explanation format does not matter; but explanations do -- An Eggsbert study on explaining Bayesian Optimisation tasks
- A Giant-Step Baby-Step Classifier For Scalable and Real-Time Anomaly Detection In Industrial Control Systems and Water Treatment Systems
- Explanations Go Linear: Post-hoc Explainability for Tabular Data with Interpretable Meta-Encoding
- Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
- ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
- BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems
- Interpretable Context Methodology: Folder Structure as Agentic Architecture
- TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models
- GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
- AI Supply Chains: An Emerging Ecosystem of AI Actors, Products, and Services
- ODExAI: A Comprehensive Object Detection Explainable AI Evaluation
- Newton-Puiseux Analysis for Interpretability and Calibration of Complex-Valued Neural Networks
- SSA-UNet: Advanced Precipitation Nowcasting via Channel Shuffling
- Two Means to an End Goal: Connecting Explainability and Contestability in the Regulation of Public Sector AI
- Designing KRIYA: An AI Companion for Wellbeing Self-Reflection
- LIME-LLM: Probing Models with Fluent Counterfactuals, Not Broken Text
- An Explainable Market Integrity Monitoring System with Multi-Source Attention Signals and Transparent Scoring
- Forecasting Equity Correlations with Hybrid Transformer Graph Neural Network
- Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts
- Device-Native Autonomous Agents for Privacy-Preserving Negotiations
- CognitionNet: A Collaborative Neural Network for Play Style Discovery in Online Skill Gaming Platform
- Exploring Local Interpretable Model-Agnostic Explanations for Speech Emotion Recognition with Distribution-Shift
- Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution
- Bridging Econometrics and AI: VaR Estimation via Reinforcement Learning and GARCH Models
- Splits! A Flexible Dataset and Evaluation Framework for Sociocultural Linguistic Investigation
- What Makes for a Good Saliency Map? Comparing Strategies for Evaluating Saliency Maps in Explainable AI (XAI)
- Learning Explainable Dense Reward Shapes via Bayesian Optimization
- Intrinsic Barriers to Explaining Deep Foundation Models
- Surrogate Fitness Metrics for Interpretable Reinforcement Learning
- Here Comes the Explanation: A Shapley Perspective on Multi-contrast Medical Image Segmentation
- A Comparative Study of Explainable AI Methods: Model-Agnostic vs. Model-Specific Approaches
- Task-based Loss Functions in Computer Vision: A Comprehensive Review
- Interpretable machine learning for imbalanced pedestrian injury severity prediction in urban Jordan
- Explainability for Embedding AI: Aspirations and Actuality
- Mathematical Programming Models for Exact and Interpretable Formulation of Neural Networks
- Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
- Leakage and Interpretability in Concept-Based Models
- Transformation of audio embeddings into interpretable, concept-based representations
- Learning to Attribute with Attention
- Long-context Non-factoid Question Answering in Indic Languages
- Enhancing Multilingual Sentiment Analysis with Explainability for Sinhala, English, and Code-Mixed Content
- Data science vs. statistics: two cultures?
- Machine learning for integrating data in biology and medicine: Principles, practice, and opportunities
- Deep convolutions for in-depth automated rock typing
- Decoding Vision Transformers: the Diffusion Steering Lens
- A Reinforcement Learning Method to Factual and Counterfactual Explanations for Session-based Recommendation
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
- PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
- Readable Twins of Unreadable Models
- Representation Learning for Tabular Data: A Comprehensive Survey
- Cross-Sectional Heterogeneity in LSTM Networks for Financial Time Series
- Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation
- Evidential Rule Learning for Interpretable Classification with Abstention
- IMMENSE: Inductive Multi-perspective User Classification in Social Networks
- Velocity- and Regime-Aware Detection of Intraday Options Market Manipulation, with Explainable Attribution
- Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support
- eXplainable AI for data driven control: an inverse optimal control approach
- Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
- Quantum Phases Classification Using Quantum Machine Learning with SHAP-Driven Feature Selection
- GlyTwin: Digital Twin for Glucose Control in Type 1 Diabetes Through Optimal Behavioral Modifications Using Patient-Centric Counterfactuals
- Explainability and Continual Learning meet Federated Learning at the Network Edge
- Are We Merely Justifying Results ex Post Facto? Quantifying Explanatory Inversion in Post-Hoc Model Explanations
- Beyond Black-Box Predictions: Identifying Marginal Feature Effects in Tabular Transformer Networks
- A constraints-based approach to fully interpretable neural networks for detecting learner behaviors
- Evaluating the robustness of explainable AI in medical image recognition under natural and adversarial data corruption
- ColonScopeX: Leveraging Explainable Expert Systems with Multimodal Data for Improved Early Diagnosis of Colorectal Cancer
- Trustworthy AI Must Account for Interactions
- Beware of "Explanations" of AI
- Outcome-Based Education: Evaluating Students' Perspectives Using Transformer
- From ChatGPT to DeepSeek AI: A Comprehensive Analysis of Evolution, Deviation, and Future Implications in AI-Language Models
- Unlocking Neural Transparency: Jacobian Maps for Explainable AI in Alzheimer's Detection
- Interpretable Multimodal Learning for Tumor Protein-Metal Binding: Progress, Challenges, and Perspectives
- From Questions to Insights: Exploring XAI Challenges Reported on Stack Overflow Questions
- Am I Being Treated Fairly? A Conceptual Framework for Individuals to Ascertain Fairness
Discussions
- “Why Should I Trust You?” Explaining the Predictions of Any Classifier [pdf] [hn, 3 points, 0 comments]
- [pdf] Explaining the Predictions of Any Classifier [hn, 1 points, 0 comments]
- Explaining black box models with LIME [hn, 1 points, 0 comments]
- “Why Should I Trust You?”: Explaining the Predictions of Any Classifier [hn, 1 points, 1 comments]
- Bringing trust into machine learning http://arxiv.org/abs/1602.04938 [bsky, 0 points, 0 comments]
Related