Ayonrinde, Kola
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
2025/03/12 by Adam Karvonen, Karvonen, Adam, Can Rager +25 · 42 citations
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
2024/10/15 by Kola Ayonrinde, Michael Pearce, Ayonrinde, Kola +3 · 13 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Information Theory (cs.IT) #Machine Learning (cs.LG)
- Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
2024/11/04 by Ayonrinde, Kola · 4 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
2025/05/01 by Kola Ayonrinde, Ayonrinde, Kola, Louis Jaburi +1 · 5 citations
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Philosophy and History of Science #Statistical and Computational Modeling
- Auditing Games for Sandbagging
2025/12/08 by Jordan Taylor, Sid Black, Taylor, Jordan +23 · 1 voice · 2 citations
Computer Science · #Advanced Malware Detection Techniques #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #FOS: Computer and information sciences #Software Engineering Research #cs.AI