Explanation in Artificial Intelligence: Insights from the Social\n Sciences
2017/06/22 by Tim Miller, Miller, Tim · 126 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning in Healthcare
paper · pdf · doi:10.48550/arxiv.1706.07269
openalex publication_date 2017/06/22 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
Abstract
There has been a recent resurgence in the area of explainable artificial\nintelligence as researchers and practitioners seek to make their algorithms\nmore understandable. Much of this research is focused on explicitly explaining\ndecisions or actions to a human observer, and it should not be controversial to\nsay that looking at how humans explain to each other can serve as a useful\nstarting point for explanation in artificial intelligence. However, it is fair\nto say that most work in explainable artificial intelligence uses only the\nresearchers' intuition of what constitutes a `good' explanation. There exists\nvast and valuable bodies of research in philosophy, psychology, and cognitive\nscience of how people define, generate, select, evaluate, and present\nexplanations, which argues that people employ certain cognitive biases and\nsocial expectations towards the explanation process. This paper argues that the\nfield of explainable artificial intelligence should build on this existing\nresearch, and reviews relevant papers from philosophy, cognitive\npsychology/science, and social psychology, which study these topics. It draws\nout some important findings, and discusses ways that these can be infused with\nwork on explainable artificial intelligence.\n
Cited by
- Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
- From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation
- Coherent without Grounding, Grounded without Success: The Bidirectional Coherence Paradox in Artificial Epistemic Agents
- Cross-modal Counterfactual Explanations: Uncovering Decision Factors and Dataset Biases in Subjective Classification
- Teaching and Critiquing Conceptualization and Operationalization in NLP
- Unifying Causal Reinforcement Learning: Survey, Taxonomy, Algorithms and Applications
- The Agony of Opacity: Foundations for Reflective Interpretability in AI-Mediated Mental Health Support
- The Influence of Human-like Appearance on Expected Robot Explanations
- STACHE: Local Black-Box Explanations for Reinforcement Learning Policies
- ContextualSHAP : Enhancing SHAP Explanations Through Contextual Language Generation
- Financial Fraud Identification and Interpretability Study for Listed Companies Based on Convolutional Neural Network
- Beyond Satisfaction: From Placebic to Actionable Explanations For Enhanced Understandability
- Human Cognitive Biases in Explanation-Based Interaction: The Case of Within and Between Session Order Effect
- A Framework for Causal Concept-based Model Explanations
- Stress-Testing Causal Claims via Cardinality Repairs
- Optimal Comprehensible Targeting
- Beyond the Black Box: A Cognitive Architecture for Explainable and Aligned AI
- Actionable and diverse counterfactual explanations incorporating domain knowledge and causal constraints
- Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
- Formal Abductive Latent Explanations for Prototype-Based Networks
- Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for Explanations
- MACIE: Multi-Agent Causal Intelligence Explainer for Collective Behavior Understanding
- llmSHAP: A Principled Approach to LLM Explainability
- Unlocking the Black Box: A Five-Dimensional Framework for Evaluating Explainable AI in Credit Risk
- T-FIX: Text-Based Explanations with Features Interpretable to eXperts
- Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
- Retrofitters, pragmatists and activists: Public interest litigation for accountable automated decision-making
- Fair and Explainable Credit-Scoring under Concept Drift: Adaptive Explanation Frameworks for Evolving Populations
- Interpretable Model-Aware Counterfactual Explanations for Random Forest
- Stop Saying "AI"
- Robustness and trustworthiness in AI: a no-go result from formal epistemology
- Imaginative Thought
- What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
- Explainability Requirements as Hyperproperties
- The seven roles of generative AI: Potential & pitfalls in combatting misinformation
- Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
- Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges
- A Multi-level Analysis of Factors Associated with Student Performance: A Machine Learning Approach to the SAEB Microdata
- Towards the Formalization of a Trustworthy AI for Mining Interpretable Models explOiting Sophisticated Algorithms
- Human-Centered LLM-Agent System for Detecting Anomalous Digital Asset Transactions
- Design Considerations for Human Oversight of AI: Insights from Co-Design Workshops and Work Design Theory
- Leveraging Association Rules for Better Predictions and Better Explanations
- Discrimination, intelligence artificielle et decisions algorithmiques
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Preliminary Quantitative Study on Explainability and Trust in AI Systems
- Discrimination, artificial intelligence, and algorithmic decision-making
- On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
- Argumentation-Based Explainability for Legal AI: Comparative and Regulatory Perspectives
- ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
- Extended Triangular Method: A Generalized Algorithm for Contradiction Separation Based Automated Deduction
- Assessing Policy Updates: Toward Trust-Preserving Intelligent User Interfaces
- From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening
- Training Feature Attribution for Vision Models
- Towards Meaningful Transparency in Civic AI Systems
- "Sometimes You Need Facts, and Sometimes a Hug": Understanding Older Adults' Preferences for Explanations in LLM-Based Conversational AI Systems
- Cluster Paths: Navigating Interpretability in Neural Networks
- Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
- Trust in Transparency: How Explainable AI Shapes User Perceptions
- Reproducibility Study of "XRec: Large Language Models for Explainable Recommendation"
- Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
- Kantian-Utilitarian XAI: Meta-Explained
- Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
- Onto-Epistemological Analysis of AI Explanations
- From Facts to Foils: Designing and Evaluating Counterfactual Explanations for Smart Environments
- Human-Centered Evaluation of RAG outputs: a framework and questionnaire for human-AI collaboration
- The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models
- Not All Explanations are Created Equal: Investigating the Pitfalls of Current XAI Evaluation
- Contrastive Concept Importance: Explaining Pairwise Class Decisions Through Automatically Extracted Concept Representations
- Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
- Toward Human-Centered Explainability: Natural Language Explanations for Anomaly Detection
- La inteligencia artificial explicable y su papel clave en la educación
- Efficient & Correct Predictive Equivalence for Decision Trees
- Looking in the mirror: A faithful counterfactual explanation method for interpreting deep image classification models
- Towards a Transparent and Interpretable AI Model for Medical Image Classifications
- Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents
- Explainability Needs in Agriculture: Exploring Dairy Farmers' User Personas
- Fairness-Aware and Interpretable Policy Learning
- Secure human oversight of AI: Threat modeling in a socio-technical context
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
- Explaining Tournament Solutions with Minimal Supports
- An Interpretable Deep Learning Model for General Insurance Pricing
- Temporal Counterfactual Explanations of Behaviour Tree Decisions
- Explainable AI in Deep Learning-Based Prediction of Solar Storms
- Triadic Fusion of Cognitive, Functional, and Causal Dimensions for Explainable LLMs: The TAXAL Framework
- TalkToAgent: A Human-centric Explanation of Reinforcement Learning Agents with Large Language Models
- An Information-Flow Perspective on Explainability Requirements: Specification and Verification
- Can AI systems have free will?
- LLM-Generated Explanations Do Not Suffice for Ultra-Strong Machine Learning
- Model Science: getting serious about verification, explanation and control of AI systems
- Interestingness First Classifiers
- From Checking to Sensemaking: A Caregiver-in-the-Loop Framework for AI-Assisted Task Verification in Dementia Care
- Toward an Interaction-Centered Approach to Robot Trustworthiness
- Rigorous Feature Importance Scores based on Shapley Value and Banzhaf Index
- Informative Post-Hoc Explanations Only Exist for Simple Functions
- Who Benefits from AI Explanations? Towards Accessible and Interpretable Systems
- To Explain Or Not To Explain: An Empirical Investigation Of AI-Based Recommendations On Social Media Platforms
- Beyond Technocratic XAI: The Who, What & How in Explanation Design
- De la innovación a la ética: pautas de uso de la inteligencia artificial en la función legislativa
- From Explainable to Explanatory Artificial Intelligence: Toward a New Paradigm for Human-Centered Explanations through Generative AI
- EICAP: Deep Dive in Assessment and Enhancement of Large Language Models in Emotional Intelligence through Multi-Turn Conversations
- Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems
- Overcoming Algorithm Aversion with Transparency: Can Transparent Predictions Change User Behavior?
- Evaluating User Experience in Conversational Recommender Systems: A Systematic Review Across Classical and LLM-Powered Approaches
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- An Appraisal-Based Approach to Human-Centred Explanations
- Foundations of Interpretable Models
- MetaExplainer: A Framework to Generate Multi-Type User-Centered Explanations for AI Systems
- Co-Producing AI: Toward an Augmented, Participatory Lifecycle
- Causal Identification of Sufficient, Contrastive and Complete Feature Sets in Image Classification
- Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
- PHAX: A Structured Argumentation Framework for User-Centered Explainable AI in Public Health and Biomedical Sciences
- Unifying Post-hoc Explanations of Knowledge Graph Completions
- Hybrid Causal Identification and Causal Mechanism Clustering
- Finding Uncommon Ground: A Human-Centered Model for Extrospective Explanations
- On Explaining Visual Captioning with Hybrid Markov Logic Networks
- Explainability in machine learning: a pedagogical perspective
- Complexity of Faceted Explanations in Propositional Abduction
- LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependency Reasoning
- Robust Explanations Through Uncertainty Decomposition: A Path to Trustworthier AI
- Exploiting Constraint Reasoning to Build Graphical Explanations for Mixed-Integer Linear Programming
- Explainable AI for online disinformation detection: Insights from a design science research project
- Argumentation meets matrix factorization: A dual perspective for explainable recommendations
- Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
- Why this and not that? A Logic-based Framework for Contrastive Explanations
- Searching for actual causes: Approximate algorithms with adjustable precision
Related