Neel Nanda
- External Invariants: A Cryptographic Trust Architecture for Institutional AI Inference
2024/06/17 by Andy Arditi, Arditi, Andy, Oscar Obeso +11 · 20 voices · 192 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
#cs.AI #cs.LG #stat.ML
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
2022/04/12 by Yuntao Bai, Bai, Yuntao, Andy Jones +59 · 501 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- In-context Learning and Induction Heads
2022/09/24 by Catherine Olsson, Olsson, Catherine, Nelson Elhage +49 · 3 voices · 111 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
- Progress measures for grokking via mechanistic interpretability
2023/01/12 by Neel Nanda, Nanda, Neel, Lawrence Chan +8 · 3 voices · 106 citations
Computer Science · Engineering · Neuroscience · #Advanced Memory and Neural Computing #Neural Networks and Applications #Neural dynamics and brain function #cs.AI #cs.LG
- Finding Neurons in a Haystack: Case Studies with Sparse Probing
2023/05/02 by Wes Gurnee, Neel Nanda, Gurnee, Wes +10 · 2 voices · 43 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.LG
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
2024/08/09 by Tom Lieberum, Lieberum, Tom, Senthooran Rajamanoharan +17 · 65 citations
Computer Science · #Machine Learning and Data Classification
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
2023/09/27 by Fred Zhang, Zhang, Fred, Neel Nanda +1 · 38 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computational and Text Analysis Methods #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Emergent Linear Representations in World Models of Self-Supervised Sequence Models
2023/09/02 by Neel Nanda, Andrew Lee, Nanda, Neel +3 · 38 citations
Computer Science · Physics and Astronomy · #Neural Networks and Reservoir Computing #Neural Networks and Applications #Model Reduction and Neural Networks
- Transcoders Find Interpretable LLM Feature Circuits
2024/06/17 by Jacob Dunefsky, Dunefsky, Jacob, Philippe Chlenski +3 · 42 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
- How to use and interpret activation patching
2024/04/23 by Stefan Heimersheim, Neel Nanda, Heimersheim, Stefan +1 · 34 citations
Computer Science · Business, Management and Accounting · #Usability and User Interface Design #Business Process Modeling and Analysis #Intelligent Tutoring Systems and Adaptive Learning
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
2024/11/21 by Javier Ferrando, Ferrando, Javier, Oscar Obeso +5 · 3 voices · 29 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Topic Modeling #cs.AI #cs.CL #cs.LG
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
2025/03/11 by Iván Arcuschin, Jett Janiak, Arcuschin, Iván +9 · 45 citations
Computer Science · Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Embodied and Extended Cognition #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
2023/07/18 by Tom Lieberum, Matthew Rahtz, Lieberum, Tom +11 · 17 citations
Computer Science · Engineering · Materials Science · #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Machine Learning in Materials Science #Topic Modeling
- Open Problems in Mechanistic Interpretability
2025/01/27 by Lee Sharkey, Bilal Chughtai, Sharkey, Lee +55 · 35 citations
Computer Science · #Natural Language Processing Techniques #Statistical and Computational Modeling
- Linear Representations of Sentiment in Large Language Models
2023/10/23 by Curt Tigges, Oskar John Hollinsworth, Tigges, Curt +5 · 16 citations
Computer Science · #Topic Modeling #Sentiment Analysis and Opinion Mining #Natural Language Processing Techniques
- Learning Multi-Level Features with Matryoshka Sparse Autoencoders
2025/03/21 by Bart Bussmann, Noa Nabeshima, Bussmann, Bart +5 · 1 voice · 29 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.LG
- Confidence Regulation Neurons in Language Models
2024/06/24 by Alessandro Stolfo, Ben Wu, Stolfo, Alessandro +11 · 18 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- BatchTopK Sparse Autoencoders
2024/12/09 by Bart Bussmann, Patrick Leask, Bussmann, Bart +3 · 22 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis
- AtP*: An efficient and scalable method for localizing LLM behaviour to components
2024/03/01 by János Kramár, Tom Lieberum, Kramár, János +5 · 14 citations
Computer Science · #Natural Language Processing Techniques #Digital Rights Management and Security
- Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
2025/04/03 by Julian Minder, Clément Dumas, Minder, Julian +7 · 3 voices · 5 citations
#cs.LG #cs.AI #cs.CL
- Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
2025/02/23 by Subhash Kantamneni, Kantamneni, Subhash, Joshua Engels +7 · 20 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
- Model Organisms for Emergent Misalignment
2025/06/13 by Edward Turner, Turner, Edward, Anna Soligo +7 · 3 voices · 17 citations
Biochemistry, Genetics and Molecular Biology · #Evolution and Genetic Dynamics #Gene Regulatory Network Analysis #cs.AI #cs.LG
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
2024/05/14 by Aleksandar Makelov, George Lange, Makelov, Aleksandar +3 · 11 citations
Computer Science · #Speech Recognition and Synthesis
- Sparse Autoencoders Do Not Find Canonical Units of Analysis
2025/02/07 by Patrick Leask, Bart Bussmann, Leask, Patrick +13 · 15 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Neural Networks and Applications
- Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
2024/02/11 by Bilal Chughtai, Chughtai, Bilal, Alan Cooney +3 · 7 citations
Economics, Econometrics and Finance · Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences #Law, Economics, and Judicial Systems #Machine Learning (cs.LG)
- Interpreting Attention Layer Outputs with Sparse Autoencoders
2024/06/25 by Connor Kissane, Kissane, Connor, Robert Krzyzanowski +7 · 7 citations
Engineering · Materials Science · Neuroscience · #Ferroelectric and Negative Capacitance Devices #Machine Learning in Materials Science #Functional Brain Connectivity Studies
- Because we have LLMs, we Can and Should Pursue Agentic Interpretability
2025/06/13 by Been Kim, John Hewitt, Kim, Been +7 · 1 voice · 2 citations
Computer Science · Social Sciences · #Multi-Agent Systems and Negotiation #Natural Language Processing Techniques #European and International Law Studies
- Base Models Know How to Reason, Thinking Models Learn When
2025/10/08 by Constantin Venhoff, Venhoff, Constantin, Iván Arcuschin +7 · 2 voices · 9 citations
#cs.AI #cs.LG
- Convergent Linear Representations of Emergent Misalignment
2025/06/13 by Anna Soligo, Edward Turner, Soligo, Anna +5 · 1 voice · 8 citations
Computer Science · Engineering · #Evolutionary Algorithms and Applications #Modular Robots and Swarm Intelligence #cs.AI #cs.LG
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
2025/07/22 by Helena Casademunt, Casademunt, Helena, C. Hsein Juang +9 · 9 citations
Computer Science · #Time Series Analysis and Forecasting #Advanced Data Compression Techniques #Gaussian Processes and Bayesian Inference
- Reasoning-Finetuning Repurposes Latent Representations in Base Models
2025/07/16 by Jake Ward, Chuqiao Lin, Ward, Jake +5 · 6 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
- Towards eliciting latent knowledge from LLMs with mechanistic interpretability
2025/05/20 by Bartosz Cywiński, Emil Ryd, Cywiński, Bartosz +5 · 4 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
- Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
2024/11/28 by Adam Karvonen, Can Rager, Karvonen, Adam +5 · 2 citations
Computer Science · #Machine Learning and Data Classification #Topic Modeling