vix.ing · top · new · best · stats · spec

Geiger, Atticus

  1. ReFT: Representation Finetuning for Language Models
    2024/04/04 by Zhengxuan Wu, Aryaman Arora, Wu, Zhengxuan +11 · 1 voice · 24 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  2. Dynabench: Rethinking Benchmarking in NLP
    2021/04/07 by Kiela, Douwe, Bartolo, Max, Nie, Yixin +16 · 37 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  3. Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
    2023/01/11 by Atticus Geiger, Geiger, Atticus, Ibeling, Duligur +9 · 34 citations
    Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences
  4. Causal Abstractions of Neural Networks
    2021/06/06 by Atticus Geiger, Geiger, Atticus, Hanson Lu +5 · 26 citations
    Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  5. AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
    2025/01/28 by Zhengxuan Wu, Aryaman Arora, Wu, Zhengxuan +13 · 1 voice · 38 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #I.2.7 #Machine Learning (cs.LG)
  6. Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
    2023/03/05 by Atticus Geiger, Geiger, Atticus, Zhengxuan Wu +7 · 18 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning #Bayesian Modeling and Causal Inference
  7. Open Problems in Mechanistic Interpretability
    2025/01/27 by Lee Sharkey, Bilal Chughtai, Sharkey, Lee +55 · 35 citations
    Computer Science · #Natural Language Processing Techniques #Statistical and Computational Modeling
  8. Linear Representations of Sentiment in Large Language Models
    2023/10/23 by Curt Tigges, Tigges, Curt, Oskar John Hollinsworth +5 · 16 citations
    Computer Science · #Topic Modeling #Sentiment Analysis and Opinion Mining #Natural Language Processing Techniques
  9. RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
    2024/02/27 by Jing Huang, Zhengxuan Wu, Huang, Jing +7 · 12 citations
    Computer Science · #Natural Language Processing Techniques
  10. Rigorously Assessing Natural Language Explanations of Neurons
    2023/09/19 by Jing Huang, Atticus Geiger, Huang, Jing +7 · 2 voices · 6 citations
    #cs.CL
  11. Inducing Causal Structure for Interpretable Neural Networks
    2021/12/01 by Atticus Geiger, Geiger, Atticus, Zhengxuan Wu +13 · 5 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Topic Modeling #Bayesian Modeling and Causal Inference
  12. Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
    2023/05/15 by Zhengxuan Wu, Wu, Zhengxuan, Atticus Geiger +6 · 7 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Topic Modeling
  13. Language Models use Lookbacks to Track Beliefs
    2025/05/20 by Nikhil Prakash, Prakash, Nikhil, Natalie Shapira +13 · 4 voices · 9 citations
    #cs.CL
  14. pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
    2024/03/12 by Zhengxuan Wu, Atticus Geiger, Wu, Zhengxuan +13 · 7 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software System Performance and Reliability
  15. Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
    2024/08/20 by Róbert Csordás, Christopher Potts, Csordás, Róbert +5 · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Image Processing and 3D Reconstruction #Machine Learning (cs.LG) #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)
  16. Enhancing Automated Interpretability with Output-Centric Feature Descriptions
    2025/01/14 by Yoav Gur-Arieh, Roy Mayan, Gur-Arieh, Yoav +7 · 6 citations
    Computer Science · #Natural Language Processing Techniques #Machine Learning and Data Classification #Topic Modeling
  17. DynaSent: A Dynamic Benchmark for Sentiment Analysis
    2020/12/30 by Potts, Christopher, Wu, Zhengxuan, Geiger, Atticus +1 · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  18. Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
    2024/09/05 by Chaudhary, Maheep, Geiger, Atticus · 5 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
  19. CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model Behavior
    2022/05/27 by Abraham, Eldar David, D'Oosterlinck, Karel, Feder, Amir +5 · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  20. Causal Abstraction with Soft Interventions
    2022/11/22 by Massidda, Riccardo, Geiger, Atticus, Icard, Thomas +1 · 2 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
  21. MIB: A Mechanistic Interpretability Benchmark
    2025/04/17 by Aaron Mueller, Atticus Geiger, Mueller, Aaron +45 · 1 voice · 6 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
  22. A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
    2024/01/23 by Zhengxuan Wu, Wu, Zhengxuan, Atticus Geiger +11 · 1 voice · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
  23. Stress-Testing Neural Models of Natural Language Inference with Multiply-Quantified Sentences
    2018/10/30 by Geiger, Atticus, Cases, Ignacio, Karttunen, Lauri +1 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  24. How Do Transformers Learn Variable Binding in Symbolic Programs?
    2025/05/27 by Yiwei Wu, Wu, Yiwei, Atticus Geiger +3 · 2 voices · 3 citations
    Computer Science · #Artificial Intelligence in Games #Evolutionary Algorithms and Applications #Metaheuristic Optimization Algorithms Research #cs.AI #cs.CL #cs.LG
  25. Relational reasoning and generalization using non-symbolic neural networks
    2020/06/14 by Geiger, Atticus, Carstensen, Alexandra, Frank, Michael C. +1 · 1 citation
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  26. How Causal Abstraction Underpins Computational Explanation
    2025/08/15 by Atticus Geiger, Jacqueline Harding, Geiger, Atticus +3 · 4 citations
    Neuroscience · Psychology · #Embodied and Extended Cognition #Philosophy and Theoretical Science #Child and Animal Learning Development
  27. Decomposing MLP Activations into Interpretable Features via Semi-Nonnegative Matrix Factorization
    2025/06/12 by Shafran, Or, Geiger, Atticus, Geva, Mor · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  28. HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
    2025/03/13 by Jiuding Sun, Sun, Jiuding, Jing Huang +11 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)