Lau, Yeu-Tong
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
2025/03/12 by Adam Karvonen, Can Rager, Karvonen, Adam +25 · 34 citations
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Applying sparse autoencoders to unlearn knowledge in language models
2024/10/25 by Farrell, Eoin, Lau, Yeu-Tong, Arthur Conmy +1 · 21 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling