vix.ing · top · new · best · stats · spec

High-dimensional variable selection

2007/04/30 by Larry Wasserman, Kathryn Roeder · 6 citations
Computer Science · Mathematics · #Advanced Statistical Methods and Models #Bayesian Methods and Mixture Models #Statistical Methods and Inference #math.ST #msc:62J05 #msc:62J07 #stat.ML #stat.TH

paper · pdf · doi:10.1214/08-aos646

published as Annals of Statistics 2009, Vol. 37, No. 5A, 2178-2201 · Published in at http://dx.doi.org/10.1214/08-AOS646 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)

openalex publication_date 2009/07/15 · arxiv created 2009/08/20 · arxiv updated 2009/12/01 · openalex created_date 2016/06/24 · openalex updated_date 2026/08/01

Abstract

This paper explores the following question: what kind of statistical guarantees can be given when doing variable selection in high dimensional models? In particular, we look at the error rates and power of some multi-stage regression methods. In the first stage we fit a set of candidate models. In the second stage we select one model by cross-validation. In the third stage we use hypothesis testing to eliminate some variables. We refer to the first two stages as "screening" and the last stage as "cleaning." We consider three screening methods: the lasso, marginal regression, and forward stepwise regression. Our method gives consistent variable selection under certain conditions.

Citations

Cited by