2019/11/06 by Dylan Slack, Slack, Dylan, Sophie Hilgard +7 · 44 citations
Computer Science · Medicine · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
paper · pdf · doi:10.48550/arxiv.1911.02508
openalex publication_date 2019/11/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
As machine learning black boxes are increasingly being deployed in domains\nsuch as healthcare and criminal justice, there is growing emphasis on building\ntools and techniques for explaining these black boxes in an interpretable\nmanner. Such explanations are being leveraged by domain experts to diagnose\nsystematic errors and underlying biases of black boxes. In this paper, we\ndemonstrate that post hoc explanations techniques that rely on input\nperturbations, such as LIME and SHAP, are not reliable. Specifically, we\npropose a novel scaffolding technique that effectively hides the biases of any\ngiven classifier by allowing an adversarial entity to craft an arbitrary\ndesired explanation. Our approach can be used to scaffold any biased classifier\nin such a way that its predictions on the input data distribution still remain\nbiased, but the post hoc explanations of the scaffolded classifier look\ninnocuous. Using extensive evaluation with multiple real-world datasets\n(including COMPAS), we demonstrate how extremely biased (racist) classifiers\ncrafted by our framework can easily fool popular explanation techniques such as\nLIME and SHAP into generating innocuous explanations which do not reflect the\nunderlying biases.\n