2025/04/29 by Z. T. He, Zhengfu He, He, Zhengfu +14 · 1 voice · 2 citations
Computer Science · Engineering · Materials Science · #Advanced Neural Network Applications #Associative property #Autoencoder #Computation and Language (cs.CL) #FOS: Computer and information sciences #Interpretability #Low-power high-performance VLSI design #Machine Learning (cs.LG) #Machine Learning in Materials Science #Neural coding #Resampling #Sparse matrix #Transformer #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2504.20938
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/04/29 · arxiv published 2025/04/29 · arxiv updated 2025/04/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We propose Low-Rank Sparse Attention (Lorsa), a sparse replacement model of Transformer attention layers to disentangle original Multi Head Self Attention (MHSA) into individually comprehensible components. Lorsa is designed to address the challenge of attention superposition to understand attention-mediated interaction between features in different token positions. We show that Lorsa heads find cleaner and finer-grained versions of previously discovered MHSA behaviors like induction heads, successor heads and attention sink behavior (i.e., heavily attending to the first token). Lorsa and Sparse Autoencoder (SAE) are both sparse dictionary learning methods applied to different Transformer components, and lead to consistent findings in many ways. For instance, we discover a comprehensive family of arithmetic-specific Lorsa heads, each corresponding to an atomic operation in Llama-3.1-8B. Automated interpretability analysis indicates that Lorsa achieves parity with SAE in interpretability while Lorsa exhibits superior circuit discovery properties, especially for features computed collectively by multiple MHSA heads. We also conduct extensive experiments on architectural design ablation, Lorsa scaling law and error analysis.