vix.ing · top · new · best · stats · spec

Kola Ayonrinde

  1. SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
    2025/03/12 by Adam Karvonen, Can Rager, Karvonen, Adam +25 · 42 citations
    Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  2. Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
    2024/10/15 by Kola Ayonrinde, Michael Pearce, Ayonrinde, Kola +3 · 13 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Information Theory (cs.IT) #Machine Learning (cs.LG)
  3. A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
    2025/05/01 by Kola Ayonrinde, Ayonrinde, Kola, Louis Jaburi +1 · 5 citations
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Philosophy and History of Science #Statistical and Computational Modeling
  4. Auditing Games for Sandbagging
    2025/12/08 by Jordan Taylor, Taylor, Jordan, Sid Black +23 · 1 voice · 2 citations
    Computer Science · #Advanced Malware Detection Techniques #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #FOS: Computer and information sciences #Software Engineering Research #cs.AI