How to use and interpret activation patching
2024/04/23 by Stefan Heimersheim, Neel Nanda, Heimersheim, Stefan +1 · 61 citations
Computer Science · Business, Management and Accounting · #Usability and User Interface Design #Business Process Modeling and Analysis #Intelligent Tutoring Systems and Adaptive Learning
paper · pdf · doi:10.48550/arxiv.2404.15255
Abstract
Activation patching is a popular mechanistic interpretability technique, but has many subtleties regarding how it is applied and how one may interpret the results. We provide a summary of advice and best practices, based on our experience using this technique in practice. We include an overview of the different ways to apply activation patching and a discussion on how to interpret the results. We focus on what evidence patching experiments provide about circuits, and on the choice of metric and associated pitfalls.
Cited by
- Grounding latent algorithm routing in transformer reasoning
- Emergent Latent-State Computation under Stochastic Volatility
- On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
- LogicCBMs: Logic-Enhanced Concept-Based Learning
- Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis
- Mechanistic Interpretability for Transformer-based Time Series Classification
- Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- Understanding Counting Mechanisms in Large Language and Vision-Language Models
- Weight-sparse transformers have interpretable circuits
- Cognitive Maps in Language Models: A Mechanistic Analysis of Spatial Planning
- Where does an LLM begin computing an instruction?
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- Addressing divergent representations from causal interventions on neural networks
- How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
- Consistency Training Helps Stop Sycophancy and Jailbreaks
- DARTS-GT: Differentiable Architecture Search for Graph Transformers with Quantifiable Instance-Specific Interpretability Analysis
- Medical Interpretability and Knowledge Maps of Large Language Models
- The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
- How do LLMs Compute Verbal Confidence
- Toward Mechanistic Explanation of Deductive Reasoning in Language Models
- Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
- Query Circuits: Explaining How Language Models Answer User Prompts
- From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
- Causally-Enhanced Reinforcement Policy Optimization
- CLUE: Conflict-guided Localization for LLM Unlearning Framework
- Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
- Pathological Truth Bias in Vision-Language Models
- Beyond Transcription: Mechanistic Interpretability in ASR
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- A Review of Developmental Interpretability in Large Language Models
- eDIF: A European Deep Inference Fabric for Remote Interpretability of LLM
- Dissecting Persona-Driven Reasoning in Language Models via Activation Patching
- How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
- Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning
- BlueGlass: A Framework for Composite AI Safety
- LAWFUL: Law-Aligned Witness for Faithful Use of Latents
- Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
- Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
- Learning Modular Exponentiation with Transformers
- Multiple Streams of Knowledge Retrieval: Enriching and Recalling in Transformers
- Large Language Models as Psychological Simulators: A Methodological Guide
- Rethinking Explainability in the Era of Multimodal AI
- Tug-of-war between idioms' figurative and literal interpretations in LLMs
- Circuit Stability Characterizes Language Model Generalization
- Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
- Divergent large language model predictions from convergent representations in ambiguous word pairs
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
- Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
- Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates
- Are We Paying Attention to Her? Investigating Gender Disambiguation and Attention in Machine Translation
- Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
- The production of meaning in the processing of natural language
- Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories
- Evidence for feature-specific error correction in LLMs
- How Transparent is DiffusionGemma?
- Steered LLM Activations are Non-Surjective
- Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning
- Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages
- Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims
- Towards Quantifying Commonsense Reasoning with Mechanistic Insights
- How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
Related