2017/05/16 by Sen Zhao, Daniela Witten, Zhao, Sen +3 · 1 citation
Computer Science · Mathematics · #Advanced Statistical Methods and Models #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (stat.ML) #Methodology (stat.ME) #Neural Networks and Applications #Statistical Methods and Inference #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.1705.05543
openalex publication_date 2017/05/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
A great deal of interest has recently focused on conducting inference on the\nparameters in a high-dimensional linear model.\n In this paper, we consider a simple and very na "ive two-step procedure for\nthis task, in which we (i) fit a lasso model in order to obtain a subset of the\nvariables, and (ii) fit a least squares model on the lasso-selected set.\nConventional statistical wisdom tells us that we cannot make use of the\nstandard statistical inference tools for the resulting least squares model\n(such as confidence intervals and p-values), since we peeked at the data\ntwice: once in running the lasso, and again in fitting the least squares model.\nHowever, in this paper, we show that under a certain set of assumptions, with\nhigh probability, the set of variables selected by the lasso is identical to\nthe one selected by the noiseless lasso and is hence deterministic.\nConsequently, the na "ive two-step approach can yield asymptotically valid\ninference. We utilize this finding to develop the \na "ive confidence\ninterval, which can be used to draw inference on the regression coefficients\nof the model selected by the lasso, as well as the \na "ive score test,\nwhich can be used to test the hypotheses regarding the full-model regression\ncoefficients.\n