2025/06/28 by Ziyang Guo, Berk Ustun, Hullman, Jessica +3 · 2 voices · 1 citation
Computer Science · Neuroscience · #Adversarial Robustness in Machine Learning #Embodied and Extended Cognition #Explainable Artificial Intelligence (XAI) #cs.AI #stat.ML
paper · pdf · doi:10.48550/arxiv.2506.22740
openalex publication_date 2025/06/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Explanations of model behavior are commonly evaluated via proxy properties weakly tied to the purposes explanations serve in practice. We contribute a decision theoretic framework that treats explanations as information signals valued by the expected improvement they enable on a specified decision task. This approach yields three distinct estimands: 1) a theoretical benchmark that upperbounds achievable performance by any agent with the explanation, 2) a human-complementary value that quantifies the theoretically attainable value that is not already captured by a baseline human decision policy, and 3) a behavioral value representing the causal effect of providing the explanation to human decision-makers. We instantiate these definitions in a practical validation workflow, and apply them to assess explanation potential and interpret behavioral effects in human-AI decision support and mechanistic interpretability.