vix.ing · top · new · best · stats · spec

Rajamanoharan, Senthooran

  1. Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
    2025/03/11 by Iván Arcuschin, Arcuschin, Iván, Jett Janiak +9 · 1 voice · 49 citations
    Computer Science · Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Embodied and Extended Cognition #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
  2. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
    2024/08/09 by Tom Lieberum, Lieberum, Tom, Senthooran Rajamanoharan +17 · 75 citations
    Computer Science · #Machine Learning and Data Classification
  3. Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
    2024/07/19 by Rajamanoharan, Senthooran, Lieberum, Tom, Sonnerat, Nicolas +4 · 59 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  4. Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
    2024/11/21 by Javier Ferrando, Oscar Obeso, Ferrando, Javier +5 · 3 voices · 32 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Topic Modeling #cs.AI #cs.CL #cs.LG
  5. Improving Dictionary Learning with Gated Sparse Autoencoders
    2024/04/24 by Rajamanoharan, Senthooran, Conmy, Arthur, Smith, Lewis +5 · 28 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
    2025/02/23 by Subhash Kantamneni, Kantamneni, Subhash, Joshua Engels +7 · 24 citations
    Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
  7. Model Organisms for Emergent Misalignment
    2025/06/13 by Edward Turner, Anna Soligo, Turner, Edward +7 · 3 voices · 18 citations
    Biochemistry, Genetics and Molecular Biology · #Evolution and Genetic Dynamics #Gene Regulatory Network Analysis #cs.AI #cs.LG
  8. An Approach to Technical AGI Safety and Security
    2025/04/02 by Shah, Rohin, Irpan, Alex, Turner, Alexander Matt +27 · 20 citations
    #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
    2025/07/07 by Emmons, Scott, Jenner, Erik, Elson, David K. +5 · 23 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  10. Convergent Linear Representations of Emergent Misalignment
    2025/06/13 by Anna Soligo, Soligo, Anna, Edward Turner +5 · 1 voice · 8 citations
    Computer Science · Engineering · #Evolutionary Algorithms and Applications #Modular Robots and Swarm Intelligence #cs.AI #cs.LG
  11. Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
    2025/07/22 by Helena Casademunt, C. Hsein Juang, Casademunt, Helena +9 · 9 citations
    Computer Science · #Time Series Analysis and Forecasting #Advanced Data Compression Techniques #Gaussian Processes and Bayesian Inference
  12. Dense SAE Latents Are Features, Not Bugs
    2025/06/18 by Sun, Xiaoqing, Stolfo, Alessandro, Engels, Joshua +4 · 6 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  13. Towards eliciting latent knowledge from LLMs with mechanistic interpretability
    2025/05/20 by Bartosz Cywiński, Emil Ryd, Cywiński, Bartosz +5 · 4 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
  14. Simple Mechanistic Explanations for Out-Of-Context Reasoning
    2025/07/10 by Wang, Atticus, Joshua Engels, Engels, Joshua +6 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
  15. Eliciting Secret Knowledge from Language Models
    2025/10/01 by Cywiński, Bartosz, Ryd, Emil, Wang, Rowan +4 · 3 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)