vix.ing · top · new · best · stats

Explanations can be manipulated and geometry is to blame

2019/06/19 by Ann-Kathrin Dombrowski, Dombrowski, Ann-Kathrin, Maximilian Alber +11 · 34 citations
Computer Science · Decision Sciences · Mathematics · Medicine · #Artificial Intelligence in Healthcare and Education #Cryptography and Security (cs.CR) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Scientific Computing and Data Management #cs.CR #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.1906.07983

openalex publication_date 2019/06/19 · arxiv created 2019/09/25 · arxiv updated 2019/09/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Explanation methods aim to make neural networks more trustworthy and interpretable. In this paper, we demonstrate a property of explanation methods which is disconcerting for both of these purposes. Namely, we show that explanations can be manipulated arbitrarily by applying visually hardly perceptible perturbations to the input that keep the network's output approximately constant. We establish theoretically that this phenomenon can be related to certain geometrical properties of neural networks. This allows us to derive an upper bound on the susceptibility of explanations to manipulations. Based on this result, we propose effective mechanisms to enhance the robustness of explanations.

Citations

Cited by

Related