vix.ing · top · new · best · stats · spec

Adaptive Monotone Shrinkage for Regression

2015/05/07 by Zhuang Ma, Dean P. Foster, Ma, Zhuang +5
Computer Science · Mathematics · #Advanced Statistical Methods and Models #Applied mathematics #Bayes estimator #Bayesian Methods and Mixture Models #Bias of an estimator #Computer science #Consistent estimator #Efficient estimator #Estimator #FOS: Computer and information sciences #Invariant estimator #James–Stein estimator #Mathematics #Mean squared error #Methodology (stat.ME) #Minimax estimator #Minimum-variance unbiased estimator #Monotone polygon #Oracle #Shrinkage estimator #Statistical Methods and Inference #Statistics #Stein's unbiased risk estimate #stat.ME

paper · pdf · doi:10.48550/arxiv.1505.01743

Appearing in Uncertainty in Artificial Intelligence (UAI) 2014

arxiv created 2015/05/07 · openalex publication_date 2015/05/07 · arxiv updated 2015/05/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We develop an adaptive monotone shrinkage estimator for regression models with the following characteristics: i) dense coefficients with small but important effects; ii) a priori ordering that indicates the probable predictive importance of the features. We capture both properties with an empirical Bayes estimator that shrinks coefficients monotonically with respect to their anticipated importance. This estimator can be rapidly computed using a version of Pool-Adjacent-Violators algorithm. We show that the proposed monotone shrinkage approach is competitive with the class of all Bayesian estimators that share the prior information. We further observe that the estimator also minimizes Stein's unbiased risk estimate. Along with our key result that the estimator mimics the oracle Bayes rule under an order assumption, we also prove that the estimator is robust. Even without the order assumption, our estimator mimics the best performance of a large family of estimators that includes the least squares estimator, constant-λ ridge estimator, James-Stein estimator, etc. All the theoretical results are non-asymptotic. Simulation results and data analysis from a model for text processing are provided to support the theory.

Related