vix.ing · top · new · best · stats · spec

Varma, Vikrant

  1. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
    2024/08/09 by Tom Lieberum, Lieberum, Tom, Senthooran Rajamanoharan +17 · 63 citations
    Computer Science · #Machine Learning and Data Classification
  2. Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
    2024/07/19 by Rajamanoharan, Senthooran, Lieberum, Tom, Sonnerat, Nicolas +4 · 53 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  3. Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
    2022/10/04 by Rohin Shah, Vikrant Varma, Shah, Rohin +11 · 1 voice · 15 citations
    Computer Science · #Reinforcement Learning in Robotics #Software Engineering Research #Software Reliability and Analysis Research #cs.LG
  4. Improving Dictionary Learning with Gated Sparse Autoencoders
    2024/04/24 by Rajamanoharan, Senthooran, Conmy, Arthur, Smith, Lewis +5 · 22 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  5. Explaining grokking through circuit efficiency
    2023/09/05 by Varma, Vikrant, Shah, Rohin, Kenton, Zachary +2 · 9 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. An Approach to Technical AGI Safety and Security
    2025/04/02 by Shah, Rohin, Irpan, Alex, Turner, Alexander Matt +27 · 17 citations
    #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  7. Challenges with unsupervised LLM knowledge discovery
    2023/12/15 by Sebastian Farquhar, Vikrant Varma, Farquhar, Sebastian +9 · 6 citations
    Computer Science · Materials Science · #Topic Modeling #Natural Language Processing Techniques #Machine Learning in Materials Science
  8. Imitating Interactive Intelligence
    2020/12/10 by Josh Abramson, Arun Ahuja, Abramson, Josh +55 · 7 citations
    Computer Science · #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Domain Adaptation and Few-Shot Learning
  9. MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
    2025/01/22 by Sebastian Farquhar, Farquhar, Sebastian, Vikrant Varma +11 · 1 voice · 4 citations
    Computer Science · #Blockchain Technology Applications and Security #cs.AI #cs.LG