vix.ing · top · new · best · stats · spec

σ-Ridge: group regularized ridge regression via empirical Bayes noise level cross-validation

2020/10/29 by Nikolaos Ignatiadis, Ignatiadis, Nikolaos, Panagiotis Lolas +1
Biochemistry, Genetics and Molecular Biology · Mathematics · #FOS: Computer and information sciences #FOS: Mathematics #Gene expression and cancer classification #Methodology (stat.ME) #Molecular Biology Techniques and Applications #Statistical Methods and Inference #Statistics Theory (math.ST)

paper · pdf · doi:10.48550/arxiv.2010.15817

openalex publication_date 2020/10/29 · openalex created_date 2024/04/11 · openalex updated_date 2026/07/28

Abstract

Features in predictive models are not exchangeable, yet common supervised models treat them as such. Here we study ridge regression when the analyst can partition the features into K groups based on external side-information. For example, in high-throughput biology, features may represent gene expression, protein abundance or clinical data and so each feature group represents a distinct modality. The analyst's goal is to choose optimal regularization parameters λ= (λ1, \dotsc, λK) -- one for each group. In this work, we study the impact of λ on the predictive risk of group-regularized ridge regression by deriving limiting risk formulae under a high-dimensional random effects model with p\asymp n as n → ∞. Furthermore, we propose a data-driven method for choosing λ that attains the optimal asymptotic risk: The key idea is to interpret the residual noise variance σ2, as a regularization parameter to be chosen through cross-validation. An empirical Bayes construction maps the one-dimensional parameter σ to the K-dimensional vector of regularization parameters, i.e., σ↦ \widehatλ(σ). Beyond its theoretical optimality, the proposed method is practical and runs as fast as cross-validated ridge regression without feature groups (K=1).

Related