vix.ing · top · new · best · stats · spec

Evaluating Attribution Methods using White-Box LSTMs

2020/10/16 by Yiding Hao, Hao, Yiding
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification

paper · pdf · doi:10.48550/arxiv.2010.08606

openalex publication_date 2020/10/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Interpretability methods for neural networks are difficult to evaluate because we do not understand the black-box models typically used to test them. This paper proposes a framework in which interpretability methods are evaluated using manually constructed networks, which we call white-box networks, whose behavior is understood a priori. We evaluate five methods for producing attribution heatmaps by applying them to white-box LSTM classifiers for tasks based on formal languages. Although our white-box classifiers solve their tasks perfectly and transparently, we find that all five attribution methods fail to produce the expected model explanations.

Citations

Related