2022/09/13 by Varsha Pendyala, Jihye Choi, Pendyala, Varsha +1 · 1 citation
Computer Science · Mathematics · Psychology · #Adversarial Robustness in Machine Learning #Anomaly detection #Artificial intelligence #Artificial neural network #Attribution #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Data mining #Deep neural networks #Domain (mathematical analysis) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Interpretability #Machine Learning (cs.LG) #Machine Learning and Data Classification #Machine learning #Mathematics #Psychology #Relation (database) #Social psychology #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2209.05690
published in arXiv (Cornell University) (Cornell University)
arxiv created 2022/09/13 · openalex publication_date 2022/09/13 · arxiv updated 2022/09/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
The interpretability of machine learning models has been an essential area of research for the safe deployment of machine learning systems. One particular approach is to attribute model decisions to high-level concepts that humans can understand. However, such concept-based explainability for Deep Neural Networks (DNNs) has been studied mostly on image domain. In this paper, we extend TCAV, the concept attribution approach, to tabular learning, by providing an idea on how to define concepts over tabular data. On a synthetic dataset with ground-truth concept explanations and a real-world dataset, we show the validity of our method in generating interpretability results that match the human-level intuitions. On top of this, we propose a notion of fairness based on TCAV that quantifies what layer of DNN has learned representations that lead to biased predictions of the model. Also, we empirically demonstrate the relation of TCAV-based fairness to a group fairness notion, Demographic Parity.