vix.ing · top · new · best · stats · spec

Structural Inference: Interpreting Small Language Models with Susceptibilities

2025/04/25 by Garrett Baker, Baker, Garrett, George Wang +5 · 2 voices · 2 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2504.18274

Abstract

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of an observable localized on a chosen component of the network. The resulting susceptibility can be estimated efficiently with local SGLD samples and factorizes into signed, per-token contributions that serve as attribution scores. We combine these susceptibilities into a response matrix whose low-rank structure separates functional modules such as multigram and induction heads in a 3M-parameter transformer.

Cited by

Discussions

Related