2021/02/07 by Scott Cheng‐Hsin Yang, Wai Keen Vong, Yang, Scott Cheng-Hsin +7
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.2102.03919
openalex publication_date 2021/02/07 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
State-of-the-art deep-learning systems use decision rules that are\nchallenging for humans to model. Explainable AI (XAI) attempts to improve human\nunderstanding but rarely accounts for how people typically reason about\nunfamiliar agents. We propose explicitly modeling the human explainee via\nBayesian Teaching, which evaluates explanations by how much they shift\nexplainees' inferences toward a desired goal. We assess Bayesian Teaching in a\nbinary image classification task across a variety of contexts. Absent\nintervention, participants predict that the AI's classifications will match\ntheir own, but explanations generated by Bayesian Teaching improve their\nability to predict the AI's judgements by moving them away from this prior\nbelief. Bayesian Teaching further allows each case to be broken down into\nsub-examples (here saliency maps). These sub-examples complement whole examples\nby improving error detection for familiar categories, whereas whole examples\nhelp predict correct AI judgements of unfamiliar cases.\n