Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI
2019/10/22 by Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Arrieta, Alejandro Barredo +21 · 241 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #cs.AI #cs.LG #cs.NE
paper · pdf · doi:10.48550/arxiv.1910.10045
67 pages, 13 figures, accepted for its publication in Information Fusion
openalex publication_date 2019/10/22 · arxiv created 2019/12/26 · arxiv updated 2019/12/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
In the last years, Artificial Intelligence (AI) has achieved a notable momentum that may deliver the best of expectations over many application sectors across the field. For this to occur, the entire community stands in front of the barrier of explainability, an inherent problem of AI techniques brought by sub-symbolism (e.g. ensembles or Deep Neural Networks) that were not present in the last hype of AI. Paradigms underlying this problem fall within the so-called eXplainable AI (XAI) field, which is acknowledged as a crucial feature for the practical deployment of AI models. This overview examines the existing literature in the field of XAI, including a prospect toward what is yet to be reached. We summarize previous efforts to define explainability in Machine Learning, establishing a novel definition that covers prior conceptual propositions with a major focus on the audience for which explainability is sought. We then propose and discuss about a taxonomy of recent contributions related to the explainability of different Machine Learning models, including those aimed at Deep Learning methods for which a second taxonomy is built. This literature analysis serves as the background for a series of challenges faced by XAI, such as the crossroads between data fusion and explainability. Our prospects lead toward the concept of Responsible Artificial Intelligence, namely, a methodology for the large-scale implementation of AI methods in real organizations with fairness, model explainability and accountability at its core. Our ultimate goal is to provide newcomers to XAI with a reference material in order to stimulate future research advances, but also to encourage experts and professionals from other disciplines to embrace the benefits of AI in their activity sectors, without any prior bias for its lack of interpretability.
Cited by
- GRExplainer: A Universal Explanation Method for Temporal Graph Neural Networks
- Problems With Large Language Models for Learner Modelling: Why LLMs Alone Fall Short for Responsible Tutoring in K--12 Education
- EvoXplain: When Machine Learning Models Agree on Predictions but Disagree on Why -- Measuring Mechanistic Multiplicity Across Training Runs
- Beyond Adoption Intention How Trust in Augmented Analytics Relates to Perceived Decision Quality Among Non-Technical BI Users
- Explainability Methods for Hardware Trojan Detection: A Systematic Comparison
- A Model of Causal Explanation on Neural Networks for Tabular Data
- Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
- PILAR: Personalizing Augmented Reality Interactions with LLM-based Human-Centric and Trustworthy Explanations for Daily Use Cases
- The Agony of Opacity: Foundations for Reflective Interpretability in AI-Mediated Mental Health Support
- Quantum Machine Learning for Climate Modelling
- Interpretable Hypothesis-Driven Trading:A Rigorous Walk-Forward Validation Framework for Market Microstructure Signals
- Explainable Artificial Intelligence for Economic Time Series: A Comprehensive Review and a Systematic Taxonomy of Methods and Concepts
- A Fast Interpretable Fuzzy Tree Learner
- When Medical AI Explanations Help and When They Harm
- An Adaptive Multi-Layered Honeynet Architecture for Threat Behavior Analysis via Deep Learning
- ContextualSHAP : Enhancing SHAP Explanations Through Contextual Language Generation
- Improving Local Fidelity Through Sampling and Modeling Nonlinearity
- When, How Long and How Much? Interpretable Neural Networks for Time Series Regression by Learning to Mask and Aggregate
- A Framework for Causal Concept-based Model Explanations
- An Approach to Joint Hybrid Decision Making between Humans and Artificial Intelligence
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model Predictions
- Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
- SHIC-XE: Viewpoint-Invariant Explainability via Dense 2D-3D Correspondences: an Application to Equine Pain Recognition
- MATCH: Engineering Transparent and Controllable Conversational XAI Systems through Composable Building Blocks
- Physics-Informed Spiking Neural Networks via Conservative Flux Quantization
- The Directed Prediction Change - Efficient and Trustworthy Fidelity Assessment for Local Feature Attribution Methods
- Toward Secure Content-Centric Approaches for 5G-Based IoT: Advances and Emerging Trends
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- Strategic Innovation Management in the Age of Large Language Models Market Intelligence, Adaptive R&D, and Ethical Governance
- Explainable Deep Convolutional Multi-Type Anomaly Detection
- Function Based Isolation Forest (FuBIF): A Unifying Framework for Interpretable Isolation-Based Anomaly Detection
- Large Language Models for Explainable Threat Intelligence
- Uncertainties in Physics-informed Inverse Problems: The Hidden Risk in Scientific AI
- Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
- POEMS: Product of Experts for Interpretable Multi-omic Integration using Sparse Decoding
- Interpretable Model-Aware Counterfactual Explanations for Random Forest
- dtControl2+ε: Trading Optimality for Explainability in MDPs via Decision Trees
- (EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations
- TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning
- "I Use ChatGPT to Humanize My Words": Affordances and Risks of ChatGPT to Autistic Users
- Unravelling impact of comorbidities on mortality risks in CKD patients during the COVID-19 pandemic: An explainable AI-driven study
- Towards Human-AI Synergy in Requirements Engineering: A Framework and Preliminary Study
- On the use of information fusion techniques to improve information quality: Taxonomy, opportunities and challenges
- Capsule Network-Based Multimodal Fusion for Mortgage Risk Assessment from Unstructured Data Sources
- Assessing the Feasibility of Early Cancer Detection Using Routine Laboratory Data: An Evaluation of Machine Learning Approaches on an Imbalanced Dataset
- Human-Centered LLM-Agent System for Detecting Anomalous Digital Asset Transactions
- ShapeX: Shapelet-Driven Post Hoc Explanations for Time Series Classification Models
- Leveraging Association Rules for Better Predictions and Better Explanations
- A Rectification-Based Approach for Distilling Boosted Trees into Decision Trees
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Towards Automated Governance: A DSL for Human-Agent Collaboration in Software Projects
- On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
- Beyond the Brightest: A Deep Learning Approach to Identifying Major and Minor Galaxy Mergers in CANDELS at z ∼ 1
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- Restricted Receptive Fields for Face Verification
- Explainable Human-in-the-Loop Segmentation via Critic Feedback Signals
- Towards Meaningful Transparency in Civic AI Systems
- Explaining raw data complexity to improve satellite onboard processing
- "Sometimes You Need Facts, and Sometimes a Hug": Understanding Older Adults' Preferences for Explanations in LLM-Based Conversational AI Systems
- Accountability Capture: How Record-Keeping to Support AI Transparency and Accountability (Re)shapes Algorithmic Oversight
- Onto-Epistemological Analysis of AI Explanations
- CORTEX: Collaborative LLM Agents for High-Stakes Alert Triage
- Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
- An Analysis of the New EU AI Act and A Proposed Standardization Framework for Machine Learning Fairness
- An Agent-Based Framework for Automated Higher-Voice Harmony Generation
- Multiplicative-Additive Constrained Models:Toward Joint Visualization of Interactive and Independent Effects
- Semantic-Inductive Attribute Selection for Zero-Shot Learning
- Does AI Coaching Prepare us for Workplace Negotiations?
- The Case for Vibe Modeling: A Missing Step in AI-Based Trustworthy Software Development
- Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments
- FA(IR)2MA-GLVQ – A hidden-feature-bias mitigation approach for fairness in classification learning based on generalized matrix learning vector quantization
- Noise, bias and data limitations: Virtual benchmarking of species distribution models' robustness for prediction and inference
- Challenges in Synchronous & Remote Collaboration Around Visualization
- La inteligencia artificial explicable y su papel clave en la educación
- Logging Requirement for Continuous Auditing of Responsible Machine Learning-based Applications
- "I think this is fair'': Uncovering the Complexities of Stakeholder Decision-Making in AI Fairness Assessment
- Cross-Attention is Half Explanation in Speech-to-Text Models
- Towards a Transparent and Interpretable AI Model for Medical Image Classifications
- Explicit vs. Implicit Biographies: Evaluating and Adapting LLM Information Extraction on Wikidata-Derived Texts
- Explainability Needs in Agriculture: Exploring Dairy Farmers' User Personas
- Out of Distribution Detection in Self-adaptive Robots with AI-powered Digital Twins
- Explainable Counterfactual Reasoning in Depression Medication Selection at Multi-Levels (Personalized and Population)
- Explainable Unsupervised Multi-Anomaly Detection and Temporal Localization in Nuclear Times Series Data with a Dual Attention-Based Autoencoder
- Clarifying Model Transparency: Interpretability versus Explainability in Deep Learning with MNIST and IMDB Examples
- The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis
- Investigating Feature Attribution for 5G Network Intrusion Detection
- Robo-Advisors Beyond Automation: Principles and Roadmap for AI-Driven Financial Planning
- Uncertainty Awareness and Trust in Explainable AI- On Trust Calibration using Local and Global Explanations
- Explainability of CNN Based Classification Models for Acoustic Signal
- Minimal Data, Maximum Clarity: A Heuristic for Explaining Optimization
- Enhancing IoMT Security with Explainable Machine Learning: A Case Study on the CICIOMT2024 Dataset
- Temporal Image Forensics: A Review and Critical Evaluation
- Triadic Fusion of Cognitive, Functional, and Causal Dimensions for Explainable LLMs: The TAXAL Framework
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- Accountability Framework for Healthcare AI Systems: Towards Joint Accountability in Decision Making
- Explainability-Driven Dimensionality Reduction for Hyperspectral Imaging
- LLM-Generated Explanations Do Not Suffice for Ultra-Strong Machine Learning
- The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms
- Cross-Attention Multimodal Fusion for Breast Cancer Diagnosis: Integrating Mammography and Clinical Data with Explainability
- Quantile Function-Based Models for Neuroimaging Classification Using Wasserstein Regression
- On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating
- Conformalized Exceptional Model Mining: Telling Where Your Model Performs (Not) Well
- Towards LLM-generated explanations for Component-based Knowledge Graph Question Answering Systems
- From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms
- On Spectral Properties of Gradient-based Explanation Methods
- Adoption of Explainable Natural Language Processing: Perspectives from Industry and Academia on Practices and Challenges
- Beyond Technocratic XAI: The Who, What & How in Explanation Design
- StreetReaderAI: Making Street View Accessible Using Context-Aware Multimodal AI
- Can AI Explanations Make You Change Your Mind?
- Neural Logic Networks for Interpretable Classification
- Graph-Based Intrusion Detection for Edge-Cloud IoT Energy Systems
- From Explainable to Explanatory Artificial Intelligence: Toward a New Paradigm for Human-Centered Explanations through Generative AI
- Detecting and explaining postpartum depression in real-time with generative artificial intelligence
- Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems
- Holistic Explainable AI (H-XAI): Extending Transparency Beyond Developers in AI-Driven Decision Making
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- An Unsupervised Deep Explainable AI Framework for Localization of Concurrent Replay Attacks in Nuclear Reactor Signals
- Six Guidelines for Trustworthy, Ethical and Responsible Automation Design
- TeSent: A Benchmark Dataset for Fairness-aware Explainable Sentiment Classification in Telugu
- Foundations of Interpretable Models
- Co-Producing AI: Toward an Augmented, Participatory Lifecycle
- Causal Explanation of Concept Drift -- A Truly Actionable Approach
- Multi-Hazard Early Warning Systems for Agriculture with Featural-Temporal Explanations
- Unifying Post-hoc Explanations of Knowledge Graph Completions
- Finding Uncommon Ground: A Human-Centered Model for Extrospective Explanations
- Towards trustworthy AI in materials mechanics through domain-guided attention
- LargeMvC-Net: Anchor-based Deep Unfolding Network for Large-scale Multi-view Clustering
- Explainability in machine learning: a pedagogical perspective
- Advancing Responsible Innovation in Agentic AI: A study of Ethical Frameworks for Household Automation
- Enhancing IoT Intrusion Detection Systems through Adversarial Training
- Explainable AI guided unsupervised fault diagnostics for high-voltage circuit breakers
- Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack
- Faithful, Interpretable Chest X-ray Diagnosis with Anti-Aliased B-cos Networks
- Exploiting Constraint Reasoning to Build Graphical Explanations for Mixed-Integer Linear Programming
- Explainable Evidential Clustering
- OrdShap: Feature Position Importance for Sequential Black-Box Models
- Trustworthy Tree-based Machine Learning by MoS2 Flash-based Analog CAM with Inherent Soft Boundaries
- Explainable AI for online disinformation detection: Insights from a design science research project
- A Survey on Interpretability in Visual Recognition
- Survey on Methods for Detection, Classification and Location of Faults in Power Systems Using Artificial Intelligence
- Counterfactual Visual Explanation via Causally-Guided Adversarial Steering
- From Classical Machine Learning to Emerging Foundation Models: Review on Multimodal Data Integration for Cancer Research
- Interpretability-Aware Pruning for Efficient Medical Image Analysis
- A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
- Bridging the gap in FER: addressing age bias in deep learning
- Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs
- MARBLE: A Multi-Agent Rule-Based LLM Reasoning Engine for Accident Severity Prediction
- Solution Space Path Planning: A Real-Time Human-Centered Path Planning Algorithm for En-Route Air Traffic Control
- Towards integration of Privacy Enhancing Technologies in Explainable Artificial Intelligence
- Quantum Stochastic Walks for Portfolio Optimization: Theory and Implementation on Financial Networks
- Human-Centered Explainability in Interactive Information Systems: A Survey
- A versatile XAI-based framework for efficient and explainable intrusion detection systems
- Interpretable by Design: MH-AutoML for Transparent and Efficient Android Malware Detection without Compromising Performance
- The Attribution Crisis in LLM Search Results
- IXAII: An Interactive Explainable Artificial Intelligence Interface for Decision Support Systems
- Gradient-Based Neuroplastic Adaptation for Concurrent Optimization of Neuro-Fuzzy Networks
- Understanding plant phenotypes in crop breeding through explainable AI
- Explainable AI for Radar Resource Management: Modified LIME in Deep Reinforcement Learning
- Autonomous Cyber Resilience via a Co-Evolutionary Arms Race within a Fortified Digital Twin Sandbox
- Toward the Explainability of Protein Language Models
- Interpretable Hybrid Machine Learning Models Using FOLD-R++ and Answer Set Programming
- FREQuency ATTribution: Benchmarking Frequency-based Occlusion for Time Series Data
- Pathwise Explanation of ReLU Neural Networks
- The Role of Explanation Styles and Perceived Accuracy on Decision Making in Predictive Process Monitoring
- Development of a persuasive User Experience Research (UXR) Point of View for Explainable Artificial Intelligence (XAI)
- Empirically derived evaluation requirements for responsible deployments of AI in safety-critical settings
- Towards Desiderata-Driven Design of Visual Counterfactual Explainers
- Mxplainer: Explain and Learn Insights by Imitating Mahjong Agents
- Towards Explaining Monte-Carlo Tree Search by Using Its Enhancements
- Beyond Shapley Values: Cooperative Games for the Interpretation of Machine Learning Models
- Interpretable Classification of Levantine Ceramic Thin Sections via Neural Networks
- Explainability in Context: A Multilevel Framework Aligning AI Explanations with Stakeholder with LLMs
- Improving LLMs with a knowledge from databases
- Interpretable Few-Shot Image Classification via Prototypical Concept-Guided Mixture of LoRA Experts
- TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- Causal Explanations Over Time: Articulated Reasoning for Interactive Environments
- Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
- Composable Building Blocks for Controllable and Transparent Interactive AI Systems
- Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
- Co-constructing Explanations for AI Systems using Provenance
- Do ATCOs Need Explanations, and Why? Towards ATCO-Centered Explainable AI for Conflict Resolution Advisories
- Efficient Preimage Approximation for Neural Network Certification
- Privacy Risks and Preservation Methods in Explainable Artificial Intelligence: A Scoping Review
- UGCE: User-Guided Incremental Counterfactual Exploration
- A Feature-level Bias Evaluation Framework for Facial Expression Recognition Models
- Exact Expansion Formalism for Transport Properties of Heterogeneous Materials Characterized by Arbitrary Continuous Random Fields
- The Bidirectional Relationship Between XAI and Regulation: Operationalizing XAI for the AI Act
- Unveil Sources of Uncertainty: Feature Contribution to Conformal Prediction Intervals
- Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI
- Model Discovery with Grammatical Evolution. An Experiment with Prime Numbers
- Opacity as a Feature, Not a Flaw: The LoBOX Governance Ethic for Role-Sensitive Explainability and Institutional Trust in AI
- Growable and Interpretable Neural Control with Online Continual Learning for Autonomous Lifelong Locomotion Learning Machines
- The Accountability Paradox: How Platform API Restrictions Undermine AI Transparency Mandates
- X-Edit: Detecting and Localizing Edits in Images Altered by Text-Guided Diffusion Models
- Artificial Intelligence and Modeling & Simulation: An Overview
- Financial Fraud Detection Using Explainable AI and Stacking Ensemble Methods
- Discovering Concept Directions from Diffusion-based Counterfactuals via Latent Clustering
- Navigating the Rashomon Effect: How Personalization Can Help Adjust Interpretable Machine Learning Models to Individual Users
- See What I Mean? CUE: A Cognitive Model of Understanding Explanations
- Wasserstein Distances Made Explainable: Insights into Dataset Shifts and Transport Phenomena
- PointExplainer: Towards Transparent Parkinson's Disease Diagnosis
- Artificial Intelligence in Government: Why People Feel They Lose Control
- Enhancing ML Model Interpretability: Leveraging Fine-Tuned Large Language Models for Better Understanding of AI
- Dynamic Forecasting and Temporal Feature Evolution of Stock Repurchases in Listed Companies Using Attention-Based Deep Temporal Networks
- Artificial intelligence and downscaling global climate model future projections
- Embodied Explainability and Ontological Obstacles: Why We Struggle to Explain the Answers of Large Language Models (LLMs)
- Explanations as Bias Detectors: A Critical Study of Local Post-hoc XAI Methods for Fairness Exploration
- Explainability and justification of automatic-decision making: A conceptual framework and a practical application
- Thoughts without Thinking: Reconsidering the Explanatory Value of Chain-of-Thought Reasoning in LLMs through Agentic Pipelines
- A Formal Framework for the Explanation of Finite Automata Decisions
- MonoLoss: A Training Objective for Interpretable Monosemantic Representations
- RuleKit 2: Faster and simpler rule learning
- Explanation format does not matter; but explanations do -- An Eggsbert study on explaining Bayesian Optimisation tasks
- ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
- A Design Framework for operationalizing Trustworthy Artificial Intelligence in Healthcare: Requirements, Tradeoffs and Challenges for its Clinical Adoption
- Newton-Puiseux Analysis for Interpretability and Calibration of Complex-Valued Neural Networks
- DiCE-Extended: A Robust Approach to Counterfactual Explanations in Machine Learning
- SSA-UNet: Advanced Precipitation Nowcasting via Channel Shuffling
- Reliable and efficient inverse analysis using physics-informed neural networks with normalized distance functions and adaptive weight tuning
- VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering
- A New Framework for Explainable Rare Cell Identification in Single-Cell Transcriptomics Data
- AI Forensics Across White-, Grey-, and Black-Box Access: A Process Model and Research Agenda for Post-Incident Investigation of AI Systems
- Physics-Informed Neural Networks for the safety analysis of nuclear reactors
- Testing Individual Fairness in Graph Neural Networks
- SoK: How Frontier AI Reshapes System-Level Security Risk Dynamics in Critical Infrastructure
- Do We Need Responsible XR? Drawing on Responsible AI to Inform Ethical Research and Practice into XRAI / the Metaverse
- Crisp complexity of fuzzy classifiers
- Towards responsible AI for education: Hybrid human-AI to confront the Elephant in the room
- Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability
- An overview of model uncertainty and variability in LLM-based sentiment analysis. Challenges, mitigation strategies and the role of explainability
- Readable Twins of Unreadable Models
- Evolutionary Reinforcement Learning for Interpretable Decision-Making in Supply Chain Management
- Towards Human-Centered Early Prediction Models for Academic Performance in Real-World Contexts
- In Terms of Explainability: Refining Requirements for Self-Explainable Systems
- Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
- Uncovering the Structure of Explanation Quality with Spectral Analysis
- Towards an Evaluation Framework for Explainable Artificial Intelligence Systems for Health and Well-being
- Focal Cortical Dysplasia Type II Detection Using Cross Modality Transfer Learning and Grad-CAM in 3D-CNNs for MRI Analysis
- Explaining Youth Driver Licensing Determinants Using XGBoost and SHAP
- Evaluating the robustness of explainable AI in medical image recognition under natural and adversarial data corruption
- Expectations, Explanations, and Embodiment: Attempts at Robot Failure Recovery
- A Trustworthy By Design Classification Model for Building Energy Retrofit Decision Support
- AEGIS: Human Attention-based Explainable Guidance for Intelligent Vehicle Systems
- The challenge of uncertainty quantification of large language models in medicine
Related