vix.ing · top · new · best · stats · spec

Optimal ablation for interpretability

2024/09/16 by Mei Li, Lucas Janson, Li, Maximilian +1 · 6 citations
Engineering · #FOS: Computer and information sciences #Fault Detection and Control Systems #Machine Learning (cs.LG) #Reservoir Engineering and Simulation Methods

paper · pdf · doi:10.48550/arxiv.2409.09951

openalex publication_date 2024/09/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Interpretability studies often involve tracing the flow of information through machine learning models to identify specific model components that perform relevant computations for tasks of interest. Prior work quantifies the importance of a model component on a particular task by measuring the impact of performing ablation on that component, or simulating model inference with the component disabled. We propose a new method, optimal ablation (OA), and show that OA-based component importance has theoretical and empirical advantages over measuring importance via other ablation methods. We also show that OA-based component importance can benefit several downstream interpretability tasks, including circuit discovery, localization of factual recall, and latent prediction.

Cited by

Related