2021/01/27 by Hana Šinkovec, Georg Heinze, Šinkovec, Hana +5 · 1 citation
Environmental Science · Mathematics · #Advanced Statistical Methods and Models #FOS: Computer and information sciences #Forest ecology and management #Methodology (stat.ME) #Statistical Methods and Inference
paper · pdf · doi:10.48550/arxiv.2101.11230
openalex publication_date 2021/01/27 · openalex created_date 2023/03/08 · openalex updated_date 2026/07/28
For finite samples with binary outcomes penalized logistic regression such as\nridge logistic regression (RR) has the potential of achieving smaller mean\nsquared errors (MSE) of coefficients and predictions than maximum likelihood\nestimation. There is evidence, however, that RR is sensitive to small or sparse\ndata situations, yielding poor performance in individual datasets. In this\npaper, we elaborate this issue further by performing a comprehensive simulation\nstudy, investigating the performance of RR in comparison to Firth's correction\nthat has been shown to perform well in low-dimensional settings. Performance of\nRR strongly depends on the choice of complexity parameter that is usually tuned\nby minimizing some measure of the out-of-sample prediction error or information\ncriterion. Alternatively, it may be determined according to prior assumptions\nabout true effects. As shown in our simulation and illustrated by a data\nexample, values optimized in small or sparse datasets are negatively correlated\nwith optimal values and suffer from substantial variability which translates\ninto large MSE of coefficients and large variability of calibration slopes. In\ncontrast, if the degree of shrinkage is pre-specified, accurate coefficients\nand predictions can be obtained even in non-ideal settings such as encountered\nin the context of rare outcomes or sparse predictors.\n