2008/09/01 by Daniela Witten, Daniela M. Witten, Robert Tibshirani · 22 citations
Biochemistry, Genetics and Molecular Biology · Mathematics · #Algorithm #Artificial intelligence #Bioinformatics and Genomic Networks #Biology #Computer science #Covariance #Covariance matrix #Data mining #False discovery rate #Gene #Gene expression and cancer classification #Genetics #Identification (biology) #Mathematics #Multiple comparisons problem #Noise (video) #Pattern recognition (psychology) #Principal component analysis #Sample size determination #Statistic #Statistical Methods and Inference #Statistical hypothesis testing #Statistics #Test statistic #Type I and type II errors #stat.AP
paper · pdf · doi:10.1214/08-aoas182
published in The Annals of Applied Statistics 2(3) (Institute of Mathematical Statistics) · Published in at http://dx.doi.org/10.1214/08-AOAS182 the Annals of Applied Statistics (http://www.imstat.org/aoas/) by the Institute of Mathematical Statistics (http://www.imstat.org)
openalex publication_date 2008/09/01 · arxiv created 2008/11/11 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We consider the problem of testing the significance of features in high-dimensional settings. In particular, we test for differentially-expressed genes in a microarray experiment. We wish to identify genes that are associated with some type of outcome, such as survival time or cancer type. We propose a new procedure, called Lassoed Principal Components (LPC), that builds upon existing methods and can provide a sizable improvement. For instance, in the case of two-class data, a standard (albeit simple) approach might be to compute a two-sample t-statistic for each gene. The LPC method involves projecting these conventional gene scores onto the eigenvectors of the gene expression data covariance matrix and then applying an L1 penalty in order to de-noise the resulting projections. We present a theoretical framework under which LPC is the logical choice for identifying significant genes, and we show that LPC can provide a marked reduction in false discovery rates over the conventional methods on both real and simulated data. Moreover, this flexible procedure can be applied to a variety of types of data and can be used to improve many existing methods for the identification of significant features.