Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI
2025/05/18 by Herrera, Francisco · 5 citations
#Computers and Society (cs.CY) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2505.20305
Abstract
As large language models (LLMs) are increasingly deployed in sensitive domains such as healthcare, law, and education, the demand for transparent, interpretable, and accountable AI systems becomes more urgent. Explainable AI (XAI) acts as a crucial interface between the opaque reasoning of LLMs and the diverse stakeholders who rely on their outputs in high-risk decisions. This paper presents a comprehensive reflection and survey of XAI for LLMs, framed around three guiding questions: Why is explainability essential? What technical and ethical dimensions does it entail? And how can it fulfill its role in real-world deployment?
We highlight four core dimensions central to explainability in LLMs, faithfulness, truthfulness, plausibility, and contrastivity, which together expose key design tensions and guide the development of explanation strategies that are both technically sound and contextually appropriate. The paper discusses how XAI can support epistemic clarity, regulatory compliance, and audience-specific intelligibility across stakeholder roles and decision settings.
We further examine how explainability is evaluated, alongside emerging developments in audience-sensitive XAI, mechanistic interpretability, causal reasoning, and adaptive explanation systems. Emphasizing the shift from surface-level transparency to governance-ready design, we identify critical challenges and future research directions for ensuring the responsible use of LLMs in complex societal contexts. We argue that explainability must evolve into a civic infrastructure fostering trust, enabling contestability, and aligning AI systems with institutional accountability and human-centered decision-making.
Citations
- Bridging Expertise Gaps: The Role of LLMs in Human-AI Collaboration for Cybersecurity
- Generative to Agentic AI: Survey, Conceptualization, and Challenges
- LLMs for Explainable AI: A Comprehensive Survey
- A Unified Framework with Novel Metrics for Evaluating the Effectiveness of XAI Techniques in LLMs
- Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
- Confronting verbalized uncertainty: Understanding how LLM’s verbalized uncertainty influences users in AI-assisted decision-making
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- Explainable artificial intelligence (XAI): from inherent explainability to large language models
- Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster Inference
- Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
- A Comprehensive Guide to Explainable AI: From Classical Models to LLMs
- A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
- Small Language Models: Survey, Measurements, and Insights
- Conversational Complexity for Assessing Risk in Large Language Models
- XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models
- A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
- Benchmark Data Contamination of Large Language Models: A Survey
- Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
- Massive Activations in Large Language Models
- From Understanding to Utilization: A Survey on Explainability for Large Language Models
- Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
- Causal Interpretation of Self-Attention in Pre-Trained Transformers
- Explainability for Large Language Models: A Survey
- Connecting the Dots in Trustworthy Artificial Intelligence: From AI Principles, Ethics, and Key Requirements to Responsible AI Systems and Regulation
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Three Levels of AI Transparency
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Towards Faithfully Interpretable NLP Systems: How should we define and\n evaluate faithfulness?
- Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI
- Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)
- Large Language Models and Causal Inference in Collaboration: A Survey
Cited by
Related