Nanda, Neel
- External Invariants: A Cryptographic Trust Architecture for Institutional AI Inference
2024/06/17 by Andy Arditi, Oscar Obeso, Arditi, Andy +11 · 20 voices · 193 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 67 citations
#cs.AI #cs.LG #stat.ML
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
2022/04/12 by Yuntao Bai, Andy Jones, Bai, Yuntao +59 · 502 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Open Problems in Mechanistic Interpretability
2025/01/27 by Lee Sharkey, Bilal Chughtai, Sharkey, Lee +58 · 7 voices · 35 citations
Computer Science · #Natural Language Processing Techniques #Statistical and Computational Modeling #cs.LG
- In-context Learning and Induction Heads
2022/09/24 by Catherine Olsson, Nelson Elhage, Olsson, Catherine +49 · 3 voices · 112 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
- Progress measures for grokking via mechanistic interpretability
2023/01/12 by Neel Nanda, Lawrence Chan, Nanda, Neel +8 · 3 voices · 106 citations
Computer Science · Engineering · Neuroscience · #Advanced Memory and Neural Computing #Neural Networks and Applications #Neural dynamics and brain function #cs.AI #cs.LG
- Finding Neurons in a Haystack: Case Studies with Sparse Probing
2023/05/02 by Wes Gurnee, Gurnee, Wes, Neel Nanda +10 · 2 voices · 43 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.LG
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
2024/08/09 by Tom Lieberum, Senthooran Rajamanoharan, Lieberum, Tom +17 · 65 citations
Computer Science · #Machine Learning and Data Classification
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
2023/09/27 by Fred Zhang, Neel Nanda, Zhang, Fred +1 · 38 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computational and Text Analysis Methods #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Emergent Linear Representations in World Models of Self-Supervised Sequence Models
2023/09/02 by Neel Nanda, Andrew Lee, Nanda, Neel +3 · 38 citations
Computer Science · Physics and Astronomy · #Neural Networks and Reservoir Computing #Neural Networks and Applications #Model Reduction and Neural Networks
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
2024/07/19 by Rajamanoharan, Senthooran, Lieberum, Tom, Sonnerat, Nicolas +4 · 53 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Transcoders Find Interpretable LLM Feature Circuits
2024/06/17 by Jacob Dunefsky, Philippe Chlenski, Dunefsky, Jacob +3 · 42 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
- How to use and interpret activation patching
2024/04/23 by Stefan Heimersheim, Neel Nanda, Heimersheim, Stefan +1 · 35 citations
Computer Science · Business, Management and Accounting · #Usability and User Interface Design #Business Process Modeling and Analysis #Intelligent Tutoring Systems and Adaptive Learning
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
2024/11/21 by Javier Ferrando, Oscar Obeso, Ferrando, Javier +5 · 3 voices · 30 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Topic Modeling #cs.AI #cs.CL #cs.LG
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
2025/03/11 by Iván Arcuschin, Jett Janiak, Arcuschin, Iván +9 · 46 citations
Computer Science · Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Embodied and Extended Cognition #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Improving Dictionary Learning with Gated Sparse Autoencoders
2024/04/24 by Rajamanoharan, Senthooran, Conmy, Arthur, Smith, Lewis +5 · 22 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
2023/07/18 by Tom Lieberum, Lieberum, Tom, Matthew Rahtz +11 · 17 citations
Computer Science · Engineering · Materials Science · #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Machine Learning in Materials Science #Topic Modeling
- Linear Representations of Sentiment in Large Language Models
2023/10/23 by Curt Tigges, Oskar John Hollinsworth, Tigges, Curt +5 · 16 citations
Computer Science · #Topic Modeling #Sentiment Analysis and Opinion Mining #Natural Language Processing Techniques
- Universal Neurons in GPT2 Language Models
2024/01/22 by Gurnee, Wes, Horsley, Theo, Guo, Zifan Carl +5 · 18 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations
2023/02/06 by Chughtai, Bilal, Chan, Lawrence, Nanda, Neel · 13 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Representation Theory (math.RT)
- Learning Multi-Level Features with Matryoshka Sparse Autoencoders
2025/03/21 by Bart Bussmann, Bussmann, Bart, Noa Nabeshima +5 · 1 voice · 29 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.LG
- Confidence Regulation Neurons in Language Models
2024/06/24 by Alessandro Stolfo, Ben Wu, Stolfo, Alessandro +11 · 18 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- BatchTopK Sparse Autoencoders
2024/12/09 by Bart Bussmann, Patrick Leask, Bussmann, Bart +3 · 22 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis
- AtP*: An efficient and scalable method for localizing LLM behaviour to components
2024/03/01 by János Kramár, Tom Lieberum, Kramár, János +5 · 14 citations
Computer Science · #Natural Language Processing Techniques #Digital Rights Management and Security
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
2025/03/12 by Karvonen, Adam, Rager, Can, Lin, Johnny +12 · 24 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
2025/04/03 by Julian Minder, Clément Dumas, Minder, Julian +7 · 3 voices · 5 citations
#cs.LG #cs.AI #cs.CL
- Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
2025/02/23 by Subhash Kantamneni, Joshua Engels, Kantamneni, Subhash +7 · 20 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
- Model Organisms for Emergent Misalignment
2025/06/13 by Edward Turner, Turner, Edward, Anna Soligo +7 · 3 voices · 17 citations
Biochemistry, Genetics and Molecular Biology · #Evolution and Genetic Dynamics #Gene Regulatory Network Analysis #cs.AI #cs.LG
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
2024/05/14 by Aleksandar Makelov, Makelov, Aleksandar, George Lange +3 · 11 citations
Computer Science · #Speech Recognition and Synthesis
- Thought Anchors: Which LLM Reasoning Steps Matter?
2025/06/23 by Bogdan, Paul C., Macar, Uzay, Nanda, Neel +1 · 33 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Sparse Autoencoders Do Not Find Canonical Units of Analysis
2025/02/07 by Patrick Leask, Bart Bussmann, Leask, Patrick +13 · 15 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Neural Networks and Applications
- Understanding Reasoning in Thinking Language Models via Steering Vectors
2025/06/22 by Venhoff, Constantin, Arcuschin, Iván, Torr, Philip +2 · 26 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching
2023/11/28 by Makelov, Aleksandar, Lange, Georg, Nanda, Neel · 7 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- An Approach to Technical AGI Safety and Security
2025/04/02 by Shah, Rohin, Irpan, Alex, Turner, Alexander Matt +27 · 17 citations
#Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Copy Suppression: Comprehensively Understanding an Attention Head
2023/10/06 by McDougall, Callum, Conmy, Arthur, Rushing, Cody +2 · 6 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
2024/02/11 by Bilal Chughtai, Chughtai, Bilal, Alan Cooney +3 · 7 citations
Economics, Econometrics and Finance · Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences #Law, Economics, and Judicial Systems #Machine Learning (cs.LG)
- Interpreting Attention Layer Outputs with Sparse Autoencoders
2024/06/25 by Connor Kissane, Kissane, Connor, Robert Krzyzanowski +7 · 7 citations
Engineering · Materials Science · Neuroscience · #Ferroelectric and Negative Capacitance Devices #Machine Learning in Materials Science #Functional Brain Connectivity Studies
- Because we have LLMs, we Can and Should Pursue Agentic Interpretability
2025/06/13 by Been Kim, John Hewitt, Kim, Been +7 · 1 voice · 2 citations
Computer Science · Social Sciences · #Multi-Agent Systems and Negotiation #Natural Language Processing Techniques #European and International Law Studies
- Base Models Know How to Reason, Thinking Models Learn When
2025/10/08 by Constantin Venhoff, Venhoff, Constantin, Iván Arcuschin +7 · 2 voices · 9 citations
#cs.AI #cs.LG
- Convergent Linear Representations of Emergent Misalignment
2025/06/13 by Anna Soligo, Edward Turner, Soligo, Anna +5 · 1 voice · 8 citations
Computer Science · Engineering · #Evolutionary Algorithms and Applications #Modular Robots and Swarm Intelligence #cs.AI #cs.LG
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
2025/07/22 by Helena Casademunt, C. Hsein Juang, Casademunt, Helena +9 · 9 citations
Computer Science · #Time Series Analysis and Forecasting #Advanced Data Compression Techniques #Gaussian Processes and Bayesian Inference
- Steering Evaluation-Aware Language Models to Act Like They Are Deployed
2025/10/23 by Hua, Tim Tian, Qin, Andrew, Marks, Samuel +1 · 5 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- How Visual Representations Map to Language Feature Space in Multimodal LLMs
2025/06/13 by Venhoff, Constantin, Khakzar, Ashkan, Joseph, Sonia +2 · 7 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Reasoning-Finetuning Repurposes Latent Representations in Base Models
2025/07/16 by Jake Ward, Chuqiao Lin, Ward, Jake +5 · 6 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
- Towards eliciting latent knowledge from LLMs with mechanistic interpretability
2025/05/20 by Bartosz Cywiński, Emil Ryd, Cywiński, Bartosz +5 · 4 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
- N2G: A Scalable Approach for Quantifying Interpretable Neuron Representations in Large Language Models
2023/04/22 by Foote, Alex, Nanda, Neel, Kran, Esben +2 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
2024/11/28 by Adam Karvonen, Karvonen, Adam, Can Rager +5 · 2 citations
Computer Science · #Machine Learning and Data Classification #Topic Modeling
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
2025/10/14 by Julian Minder, Clément Dumas, Minder, Julian +11 · 1 voice · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
- Real-Time Detection of Hallucinated Entities in Long-Form Generation
2025/08/26 by Obeso, Oscar, Arditi, Andy, Ferrando, Javier +3 · 5 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Scaling sparse feature circuit finding for in-context learning
2025/04/18 by Kharlapenko, Dmitrii, Shabalin, Stepan, Barez, Fazl +2 · 2 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Eliciting Secret Knowledge from Language Models
2025/10/01 by Cywiński, Bartosz, Ryd, Emil, Wang, Rowan +4 · 3 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)