vix.ing · top · new · best · stats · spec

Neel Nanda

  1. External Invariants: A Cryptographic Trust Architecture for Institutional AI Inference
    2024/06/17 by Andy Arditi, Arditi, Andy, Oscar Obeso +11 · 20 voices · 192 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  2. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
    2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
    #cs.AI #cs.LG #stat.ML
  3. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
    2022/04/12 by Yuntao Bai, Bai, Yuntao, Andy Jones +59 · 501 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  4. In-context Learning and Induction Heads
    2022/09/24 by Catherine Olsson, Olsson, Catherine, Nelson Elhage +49 · 3 voices · 111 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
  5. Progress measures for grokking via mechanistic interpretability
    2023/01/12 by Neel Nanda, Nanda, Neel, Lawrence Chan +8 · 3 voices · 106 citations
    Computer Science · Engineering · Neuroscience · #Advanced Memory and Neural Computing #Neural Networks and Applications #Neural dynamics and brain function #cs.AI #cs.LG
  6. Finding Neurons in a Haystack: Case Studies with Sparse Probing
    2023/05/02 by Wes Gurnee, Neel Nanda, Gurnee, Wes +10 · 2 voices · 43 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.LG
  7. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
    2024/08/09 by Tom Lieberum, Lieberum, Tom, Senthooran Rajamanoharan +17 · 65 citations
    Computer Science · #Machine Learning and Data Classification
  8. Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
    2023/09/27 by Fred Zhang, Zhang, Fred, Neel Nanda +1 · 38 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computational and Text Analysis Methods #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  9. Emergent Linear Representations in World Models of Self-Supervised Sequence Models
    2023/09/02 by Neel Nanda, Andrew Lee, Nanda, Neel +3 · 38 citations
    Computer Science · Physics and Astronomy · #Neural Networks and Reservoir Computing #Neural Networks and Applications #Model Reduction and Neural Networks
  10. Transcoders Find Interpretable LLM Feature Circuits
    2024/06/17 by Jacob Dunefsky, Dunefsky, Jacob, Philippe Chlenski +3 · 42 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
  11. How to use and interpret activation patching
    2024/04/23 by Stefan Heimersheim, Neel Nanda, Heimersheim, Stefan +1 · 34 citations
    Computer Science · Business, Management and Accounting · #Usability and User Interface Design #Business Process Modeling and Analysis #Intelligent Tutoring Systems and Adaptive Learning
  12. Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
    2024/11/21 by Javier Ferrando, Ferrando, Javier, Oscar Obeso +5 · 3 voices · 29 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Topic Modeling #cs.AI #cs.CL #cs.LG
  13. Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
    2025/03/11 by Iván Arcuschin, Jett Janiak, Arcuschin, Iván +9 · 45 citations
    Computer Science · Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Embodied and Extended Cognition #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  14. Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
    2023/07/18 by Tom Lieberum, Matthew Rahtz, Lieberum, Tom +11 · 17 citations
    Computer Science · Engineering · Materials Science · #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Machine Learning in Materials Science #Topic Modeling
  15. Open Problems in Mechanistic Interpretability
    2025/01/27 by Lee Sharkey, Bilal Chughtai, Sharkey, Lee +55 · 35 citations
    Computer Science · #Natural Language Processing Techniques #Statistical and Computational Modeling
  16. Linear Representations of Sentiment in Large Language Models
    2023/10/23 by Curt Tigges, Oskar John Hollinsworth, Tigges, Curt +5 · 16 citations
    Computer Science · #Topic Modeling #Sentiment Analysis and Opinion Mining #Natural Language Processing Techniques
  17. Learning Multi-Level Features with Matryoshka Sparse Autoencoders
    2025/03/21 by Bart Bussmann, Noa Nabeshima, Bussmann, Bart +5 · 1 voice · 29 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.LG
  18. Confidence Regulation Neurons in Language Models
    2024/06/24 by Alessandro Stolfo, Ben Wu, Stolfo, Alessandro +11 · 18 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  19. BatchTopK Sparse Autoencoders
    2024/12/09 by Bart Bussmann, Patrick Leask, Bussmann, Bart +3 · 22 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis
  20. AtP*: An efficient and scalable method for localizing LLM behaviour to components
    2024/03/01 by János Kramár, Tom Lieberum, Kramár, János +5 · 14 citations
    Computer Science · #Natural Language Processing Techniques #Digital Rights Management and Security
  21. Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
    2025/04/03 by Julian Minder, Clément Dumas, Minder, Julian +7 · 3 voices · 5 citations
    #cs.LG #cs.AI #cs.CL
  22. Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
    2025/02/23 by Subhash Kantamneni, Kantamneni, Subhash, Joshua Engels +7 · 20 citations
    Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
  23. Model Organisms for Emergent Misalignment
    2025/06/13 by Edward Turner, Turner, Edward, Anna Soligo +7 · 3 voices · 17 citations
    Biochemistry, Genetics and Molecular Biology · #Evolution and Genetic Dynamics #Gene Regulatory Network Analysis #cs.AI #cs.LG
  24. Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
    2024/05/14 by Aleksandar Makelov, George Lange, Makelov, Aleksandar +3 · 11 citations
    Computer Science · #Speech Recognition and Synthesis
  25. Sparse Autoencoders Do Not Find Canonical Units of Analysis
    2025/02/07 by Patrick Leask, Bart Bussmann, Leask, Patrick +13 · 15 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Neural Networks and Applications
  26. Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
    2024/02/11 by Bilal Chughtai, Chughtai, Bilal, Alan Cooney +3 · 7 citations
    Economics, Econometrics and Finance · Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences #Law, Economics, and Judicial Systems #Machine Learning (cs.LG)
  27. Interpreting Attention Layer Outputs with Sparse Autoencoders
    2024/06/25 by Connor Kissane, Kissane, Connor, Robert Krzyzanowski +7 · 7 citations
    Engineering · Materials Science · Neuroscience · #Ferroelectric and Negative Capacitance Devices #Machine Learning in Materials Science #Functional Brain Connectivity Studies
  28. Because we have LLMs, we Can and Should Pursue Agentic Interpretability
    2025/06/13 by Been Kim, John Hewitt, Kim, Been +7 · 1 voice · 2 citations
    Computer Science · Social Sciences · #Multi-Agent Systems and Negotiation #Natural Language Processing Techniques #European and International Law Studies
  29. Base Models Know How to Reason, Thinking Models Learn When
    2025/10/08 by Constantin Venhoff, Venhoff, Constantin, Iván Arcuschin +7 · 2 voices · 9 citations
    #cs.AI #cs.LG
  30. Convergent Linear Representations of Emergent Misalignment
    2025/06/13 by Anna Soligo, Edward Turner, Soligo, Anna +5 · 1 voice · 8 citations
    Computer Science · Engineering · #Evolutionary Algorithms and Applications #Modular Robots and Swarm Intelligence #cs.AI #cs.LG
  31. Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
    2025/07/22 by Helena Casademunt, Casademunt, Helena, C. Hsein Juang +9 · 9 citations
    Computer Science · #Time Series Analysis and Forecasting #Advanced Data Compression Techniques #Gaussian Processes and Bayesian Inference
  32. Reasoning-Finetuning Repurposes Latent Representations in Base Models
    2025/07/16 by Jake Ward, Chuqiao Lin, Ward, Jake +5 · 6 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
  33. Towards eliciting latent knowledge from LLMs with mechanistic interpretability
    2025/05/20 by Bartosz Cywiński, Emil Ryd, Cywiński, Bartosz +5 · 4 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
  34. Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
    2024/11/28 by Adam Karvonen, Can Rager, Karvonen, Adam +5 · 2 citations
    Computer Science · #Machine Learning and Data Classification #Topic Modeling