Galichin, Andrey
- I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
2025/03/24 by Galichin, Andrey, Dontsov, Alexey, Druzhinina, Polina +4 · 15 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
2025/09/26 by Korznikov, Anton, Galichin, Andrey, Dontsov, Alexey +3 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
2025/09/26 by Korznikov, Anton, Galichin, Andrey, Dontsov, Alexey +3 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)