Arthur Conmy
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1359 citations
#cs.CL #cs.AI
- Stealing Part of a Production Language Model
2024/03/11 by Nicholas Carlini, Carlini, Nicholas, Daniel Paleka +26 · 11 voices · 26 citations
Computer Science · #Natural Language Processing Techniques #cs.CR
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
2022/11/01 by Kevin Wang, Alexandre Variengien, Wang, Kevin +7 · 119 citations
Computer Science · Materials Science · #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning #Machine Learning in Materials Science
- Towards Automated Circuit Discovery for Mechanistic Interpretability
2023/04/28 by Arthur Conmy, Conmy, Arthur, Augustine N. Mavor-Parker +7 · 76 citations
Computer Science · Engineering · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Fault Detection and Control Systems #Machine Learning (cs.LG)
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
2024/08/09 by Tom Lieberum, Lieberum, Tom, Senthooran Rajamanoharan +17 · 63 citations
Computer Science · #Machine Learning and Data Classification
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
2025/03/11 by Iván Arcuschin, Arcuschin, Iván, Jett Janiak +9 · 45 citations
Computer Science · Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Embodied and Extended Cognition #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Attribution Patching Outperforms Automated Circuit Discovery
2023/10/16 by Aaquib Syed, Syed, Aaquib, Can Rager +3 · 20 citations
Computer Science · Materials Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Materials Science
- Open Problems in Mechanistic Interpretability
2025/01/27 by Lee Sharkey, Sharkey, Lee, Bilal Chughtai +55 · 35 citations
Computer Science · #Natural Language Processing Techniques #Statistical and Computational Modeling
- Improving Steering Vectors by Targeting Sparse Autoencoder Features
2024/11/04 by Sviatoslav Chalnev, Chalnev, Sviatoslav, Siu, Matthew +2 · 14 citations
Engineering · #Vehicle License Plate Recognition #Vehicle Dynamics and Control Systems #Autonomous Vehicle Technology and Safety
- Successor Heads: Recurring, Interpretable Attention Heads In The Wild
2023/12/14 by R. Bruce Gould, Gould, Rhys, Euan Ong +5 · 10 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Interpreting Attention Layer Outputs with Sparse Autoencoders
2024/06/25 by Connor Kissane, Kissane, Connor, Robert Krzyzanowski +7 · 7 citations
Engineering · Materials Science · Neuroscience · #Ferroelectric and Negative Capacitance Devices #Machine Learning in Materials Science #Functional Brain Connectivity Studies
- Base Models Know How to Reason, Thinking Models Learn When
2025/10/08 by Constantin Venhoff, Iván Arcuschin, Venhoff, Constantin +7 · 2 voices · 9 citations
#cs.AI #cs.LG
- How do LLMs Compute Verbal Confidence
2026/03/18 by Dharshan Kumaran, Arthur Conmy, Federico Barbero +3 · 1 voice
Computer Science · #cs.CL #cs.AI #cs.LG
- Line of Sight: On Linear Representations in VLLMs
2025/06/05 by Achyuta Rajaram, Rajaram, Achyuta, Sarah Schwettmann +5 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Topic Modeling
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
2026/07/28 by Anton de la Fuente, Arthur Conmy
Computer Science · #cs.LG