2025/09/25 by Shuofeng Zhang, Zhang, Shuofeng, Ard A. Louis +1 · 1 citation
Computer Science · Mathematics · #Bayesian Methods and Mixture Models #FOS: Computer and information sciences #FOS: Mathematics #Face and Expression Recognition #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Statistical Methods and Inference #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.2509.21181
openalex publication_date 2025/09/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
For overparameterized linear regression with isotropic Gaussian design and minimum-ℓp interpolator p∈(1,2], we give a unified, high-probability characterization for the scaling of the family of parameter norms ‖ \widehatwp ‖r r ∈ [1,p] with sample size. We solve this basic, but unresolved question through a simple dual-ray analysis, which reveals a competition between a signal *spike* and a *bulk* of null coordinates in X^\top Y, yielding closed-form predictions for (i) a data-dependent transition n_⋆ (the "elbow"), and (ii) a universal threshold r_⋆=2(p-1) that separates ‖ \widehatwp ‖r's which plateau from those that continue to grow with an explicit exponent. This unified solution resolves the scaling of *all* ℓr norms within the family r∈ [1,p] under ℓp-biased interpolation, and explains in one picture which norms saturate and which increase as n grows. We then study diagonal linear networks (DLNs) trained by gradient descent. By calibrating the initialization scale α to an effective peff(α) via the DLN separable potential, we show empirically that DLNs inherit the same elbow/threshold laws, providing a predictive bridge between explicit and implicit bias. Given that many generalization proxies depend on ‖ \widehat wp ‖r, our results suggest that their predictive power will depend sensitively on which lr norm is used.