Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
2025/10/01 by Maxime Méloux, François Portet, Méloux, Maxime +3 · 1 citation
Computer Science · Engineering · #Statistical and Computational Modeling #Explainable Artificial Intelligence (XAI) #Nuclear Engineering Thermal-Hydraulics
paper · pdf · doi:10.48550/arxiv.2510.00845
Abstract
Mechanistic Interpretability (MI) aims to reverse-engineer model behaviors by identifying functional sub-networks. Yet, the scientific validity of these findings depends on their stability. In this work, we argue that circuit discovery is not a standalone task but a statistical estimation problem built upon causal mediation analysis (CMA). We uncover a fundamental instability at this base layer: exact, single-input CMA scores exhibit high intrinsic variance, implying that the causal effect of a component is a volatile random variable rather than a fixed property. We then demonstrate that circuit discovery pipelines inherit this variance and further amplify it. Fast approximation methods, such as Edge Attribution Patching and its successors, introduce additional estimation noise, while aggregating these noisy scores over datasets leads to fragile structural estimates. Consequently, small perturbations in input data or hyperparameters yield vastly different circuits. We systematically decompose these sources of variance and advocate for more rigorous MI practices, prioritizing statistical robustness and routine reporting of stability metrics.
Citations
- The Dead Salmons of AI Interpretability
- BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
- BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
- Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
- RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching
- Think before you fit: parameter identifiability, sensitivity and uncertainty in systems biology models
- GIM: Improved Interpretability for Large Language Models
- MIB: A Mechanistic Interpretability Benchmark
- Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
- EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
- Hypothesis Testing the Circuit Hypothesis in LLMs
- The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
- Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity
- Finding Transformer Circuits with Edge Pruning
- Transcoders Find Interpretable LLM Feature Circuits
- Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking
- A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
- Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
- Attribution Patching Outperforms Automated Circuit Discovery
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
- Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark
- Physics of Language Models: Part 1, Learning Hierarchical Language Structures
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
- Towards Automated Circuit Discovery for Mechanistic Interpretability
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
- Mass-Editing Memory in a Transformer
- Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
- Locating and Editing Factual Associations in GPT
- Causal Abstractions of Neural Networks
- Refining Targeted Syntactic Evaluation of Language Models
- Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
- BLiMP: The Benchmark of Linguistic Minimal Pairs for English
- Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI
- Veridical data science
- Analyzing biological and artificial neural networks: challenges with opportunities for synergy?
- Sanity Checks for Saliency Maps
- Network Dissection: Quantifying Interpretability of Deep Visual\n Representations
- Axiomatic Attribution for Deep Networks
- Explanation in causal inference: developments in mediation and interaction
- The ASA Statement on p -Values: Context, Process, and Purpose
- Visualizing and Understanding Convolutional Networks
- Direct and Indirect Effects
- Bootstrap Methods for Standard Errors, Confidence Intervals, and Other Measures of Statistical Accuracy
Cited by
Related