vix.ing · top · new · best · stats · spec

Grounding Representation Similarity with Statistical Testing

2021/08/03 by Frances Ding, Jean-Stanislas Denain, Ding, Frances +3 · 1 citation
Computer Science · Mathematics · Psychology · #Artificial intelligence #Canonical correlation #Computer science #Data mining #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Initialization #Kernel (algebra) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Mathematics #Neural Networks and Applications #Psychology #Representation (politics) #Robustness (evolution) #Set (abstract data type) #Similarity (geometry) #Strengths and weaknesses #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.2108.01661

Accepted at NeurIPS 2021. 10 pages, 3 figures

openalex publication_date 2021/08/03 · arxiv created 2021/11/03 · arxiv updated 2021/11/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

To understand neural network behavior, recent works quantitatively compare different networks' learned representations using canonical correlation analysis (CCA), centered kernel alignment (CKA), and other dissimilarity measures. Unfortunately, these widely used measures often disagree on fundamental observations, such as whether deep networks differing only in random initialization learn similar representations. These disagreements raise the question: which, if any, of these dissimilarity measures should we believe? We provide a framework to ground this question through a concrete test: measures should have sensitivity to changes that affect functional behavior, and specificity against changes that do not. We quantify this through a variety of functional behaviors including probing accuracy and robustness to distribution shift, and examine changes such as varying random initialization and deleting principal components. We find that current metrics exhibit different weaknesses, note that a classical baseline performs surprisingly well, and highlight settings where all metrics appear to fail, thus providing a challenge set for further improvement.

Citations

Cited by

Related