vix.ing · top · new · best · stats · spec

Lau, Yeu-Tong

  1. SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
    2025/03/12 by Adam Karvonen, Can Rager, Karvonen, Adam +25 · 34 citations
    Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  2. Applying sparse autoencoders to unlearn knowledge in language models
    2024/10/25 by Farrell, Eoin, Lau, Yeu-Tong, Arthur Conmy +1 · 21 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling