vix.ing · top · new · best · stats · spec

Ayonrinde, Kola

  1. SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
    2025/03/12 by Adam Karvonen, Karvonen, Adam, Can Rager +25 · 42 citations
    Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  2. Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
    2024/10/15 by Kola Ayonrinde, Michael Pearce, Ayonrinde, Kola +3 · 13 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Information Theory (cs.IT) #Machine Learning (cs.LG)
  3. Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
    2024/11/04 by Ayonrinde, Kola · 4 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  4. A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
    2025/05/01 by Kola Ayonrinde, Ayonrinde, Kola, Louis Jaburi +1 · 5 citations
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Philosophy and History of Science #Statistical and Computational Modeling
  5. Auditing Games for Sandbagging
    2025/12/08 by Jordan Taylor, Sid Black, Taylor, Jordan +23 · 1 voice · 2 citations
    Computer Science · #Advanced Malware Detection Techniques #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #FOS: Computer and information sciences #Software Engineering Research #cs.AI