Geiger, Atticus
- ReFT: Representation Finetuning for Language Models
2024/04/04 by Zhengxuan Wu, Aryaman Arora, Wu, Zhengxuan +11 · 1 voice · 24 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Dynabench: Rethinking Benchmarking in NLP
2021/04/07 by Kiela, Douwe, Bartolo, Max, Nie, Yixin +16 · 37 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
2023/01/11 by Atticus Geiger, Geiger, Atticus, Ibeling, Duligur +9 · 34 citations
Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences
- Causal Abstractions of Neural Networks
2021/06/06 by Atticus Geiger, Geiger, Atticus, Hanson Lu +5 · 26 citations
Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
2025/01/28 by Zhengxuan Wu, Aryaman Arora, Wu, Zhengxuan +13 · 1 voice · 38 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #I.2.7 #Machine Learning (cs.LG)
- Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
2023/03/05 by Atticus Geiger, Geiger, Atticus, Zhengxuan Wu +7 · 18 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning #Bayesian Modeling and Causal Inference
- Open Problems in Mechanistic Interpretability
2025/01/27 by Lee Sharkey, Bilal Chughtai, Sharkey, Lee +55 · 35 citations
Computer Science · #Natural Language Processing Techniques #Statistical and Computational Modeling
- Linear Representations of Sentiment in Large Language Models
2023/10/23 by Curt Tigges, Tigges, Curt, Oskar John Hollinsworth +5 · 16 citations
Computer Science · #Topic Modeling #Sentiment Analysis and Opinion Mining #Natural Language Processing Techniques
- RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
2024/02/27 by Jing Huang, Zhengxuan Wu, Huang, Jing +7 · 12 citations
Computer Science · #Natural Language Processing Techniques
- Rigorously Assessing Natural Language Explanations of Neurons
2023/09/19 by Jing Huang, Atticus Geiger, Huang, Jing +7 · 2 voices · 6 citations
#cs.CL
- Inducing Causal Structure for Interpretable Neural Networks
2021/12/01 by Atticus Geiger, Geiger, Atticus, Zhengxuan Wu +13 · 5 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Topic Modeling #Bayesian Modeling and Causal Inference
- Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
2023/05/15 by Zhengxuan Wu, Wu, Zhengxuan, Atticus Geiger +6 · 7 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Topic Modeling
- Language Models use Lookbacks to Track Beliefs
2025/05/20 by Nikhil Prakash, Prakash, Nikhil, Natalie Shapira +13 · 4 voices · 9 citations
#cs.CL
- pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
2024/03/12 by Zhengxuan Wu, Atticus Geiger, Wu, Zhengxuan +13 · 7 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software System Performance and Reliability
- Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
2024/08/20 by Róbert Csordás, Christopher Potts, Csordás, Róbert +5 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Image Processing and 3D Reconstruction #Machine Learning (cs.LG) #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)
- Enhancing Automated Interpretability with Output-Centric Feature Descriptions
2025/01/14 by Yoav Gur-Arieh, Roy Mayan, Gur-Arieh, Yoav +7 · 6 citations
Computer Science · #Natural Language Processing Techniques #Machine Learning and Data Classification #Topic Modeling
- DynaSent: A Dynamic Benchmark for Sentiment Analysis
2020/12/30 by Potts, Christopher, Wu, Zhengxuan, Geiger, Atticus +1 · 2 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
2024/09/05 by Chaudhary, Maheep, Geiger, Atticus · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
- CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model Behavior
2022/05/27 by Abraham, Eldar David, D'Oosterlinck, Karel, Feder, Amir +5 · 2 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Causal Abstraction with Soft Interventions
2022/11/22 by Massidda, Riccardo, Geiger, Atticus, Icard, Thomas +1 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
- MIB: A Mechanistic Interpretability Benchmark
2025/04/17 by Aaron Mueller, Atticus Geiger, Mueller, Aaron +45 · 1 voice · 6 citations
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
- A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
2024/01/23 by Zhengxuan Wu, Wu, Zhengxuan, Atticus Geiger +11 · 1 voice · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
- Stress-Testing Neural Models of Natural Language Inference with Multiply-Quantified Sentences
2018/10/30 by Geiger, Atticus, Cases, Ignacio, Karttunen, Lauri +1 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- How Do Transformers Learn Variable Binding in Symbolic Programs?
2025/05/27 by Yiwei Wu, Wu, Yiwei, Atticus Geiger +3 · 2 voices · 3 citations
Computer Science · #Artificial Intelligence in Games #Evolutionary Algorithms and Applications #Metaheuristic Optimization Algorithms Research #cs.AI #cs.CL #cs.LG
- Relational reasoning and generalization using non-symbolic neural networks
2020/06/14 by Geiger, Atticus, Carstensen, Alexandra, Frank, Michael C. +1 · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- How Causal Abstraction Underpins Computational Explanation
2025/08/15 by Atticus Geiger, Jacqueline Harding, Geiger, Atticus +3 · 4 citations
Neuroscience · Psychology · #Embodied and Extended Cognition #Philosophy and Theoretical Science #Child and Animal Learning Development
- Decomposing MLP Activations into Interpretable Features via Semi-Nonnegative Matrix Factorization
2025/06/12 by Shafran, Or, Geiger, Atticus, Geva, Mor · 2 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
2025/03/13 by Jiuding Sun, Sun, Jiuding, Jing Huang +11 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)