vix.ing · top · new · best · stats · spec

Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation

2026/06/30 by Guo Yu, Wenlin Liu, Yulan Hu +3
Computer Science · #cs.LG

paper · pdf

Code is available at https://github.com/SydCS/OPD-Param-Analysis

arxiv created 2026/07/30 · arxiv updated 2026/07/31

Abstract

On-policy distillation (OPD) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student-generated trajectories and dense token-level teacher supervision. Yet how this hybrid training regime shapes a model remains poorly understood. We characterize the sparsity and geometry of OPD parameter updates across several language and vision-language model pairs and application settings. OPD updates are small and coordinate-sparse at checkpoint precision, while remaining distributed across layers and modules. This sparse support is operationally meaningful: masked training on the discovered subnetwork nearly recovers full-training performance. At the matrix level, the updates are numerically full-rank but spectrally concentrated. Their visible supports avoid coordinates emphasized by the source's principal structure and favor low-magnitude source coordinates, while the source singular-value spectra change little. Together, these findings show that OPD exhibits important weight-space signatures of on-policy post-training despite using dense teacher supervision.

Citations

Related