vix.ing · top · new · best · stats · spec

Better prediction by use of co-data: Adaptive group-regularized ridge\n regression

2014/11/13 by Mark A. van de Wiel, Tonje G. Lien, van de Wiel, Mark A. +7 · 1 citation
Biochemistry, Genetics and Molecular Biology · Mathematics · Medicine · #62J07 #Bioinformatics and Genomic Networks #FOS: Computer and information sciences #Gene expression and cancer classification #Methodology (stat.ME) #Radiomics and Machine Learning in Medical Imaging #Statistical Methods and Inference

paper · pdf · doi:10.48550/arxiv.1411.3496

openalex publication_date 2014/11/13 · openalex created_date 2025/10/24 · openalex updated_date 2026/07/28

Abstract

For many high-dimensional studies, additional information on the variables,\nlike (genomic) annotation or external p-values, is available. In the context of\nbinary and continuous prediction, we develop a method for adaptive\ngroup-regularized (logistic) ridge regression, which makes structural use of\nsuch 'co-data'. Here, 'groups' refer to a partition of the variables according\nto the co-data. We derive empirical Bayes estimates of group-specific\npenalties, which possess several nice properties: i) they are analytical; ii)\nthey adapt to the informativeness of the co-data for the data at hand; iii)\nonly one global penalty parameter requires tuning by cross-validation. In\naddition, the method allows use of multiple types of co-data at little extra\ncomputational effort.\n We show that the group-specific penalties may lead to a larger distinction\nbetween `near-zero' and relatively large regression parameters, which\nfacilitates post-hoc variable selection. The method, termed GRridge, is\nimplemented in an easy-to-use R-package. It is demonstrated on two cancer\ngenomics studies, which both concern the discrimination of precancerous\ncervical lesions from normal cervix tissues using methylation microarray data.\nFor both examples, GRridge clearly improves the predictive performances of\nordinary logistic ridge regression and the group lasso. In addition, we show\nthat for the second study the relatively good predictive performance is\nmaintained when selecting only 42 variables.\n

Citations

Cited by

Related