vix.ing · top · new · best · stats · spec

Exact model comparisons in the plausibility framework

2019/11/30 by Stefan Böhringer, Dietmar Lohmann
Biochemistry, Genetics and Molecular Biology · Mathematics · #Algorithm #Applied mathematics #Consistency (knowledge bases) #Data set #Discrete mathematics #Exact test #Gene expression and cancer classification #Mathematics #Multiple comparisons problem #Null hypothesis #Parametric statistics #Sample size determination #Statistical Methods and Bayesian Inference #Statistical Methods and Inference #Statistical hypothesis testing #Statistics #acm:62-04 #acm:62E15 #acm:62H10 #math.ST #msc:62-04 #msc:62E15 #msc:62H10 #stat.AP #stat.CO #stat.TH

paper · pdf · doi:10.1016/j.jspi.2021.07.013

openalex publication_date 2021/09/07 · arxiv created 2021/09/10 · arxiv updated 2021/09/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Plausibility is a formalization of exact tests for parametric models and generalizes procedures such as Fisher’s exact test. The resulting tests are based on cumulative probabilities of the probability density function and evaluate consistency with a parametric family while providing exact control of the α level for finite sample size. Model comparisons are inefficient in this approach. We generalize plausibility by incorporating weighing which allows to perform model comparisons. We show that one weighing scheme is asymptotically equivalent to the likelihood ratio test (LRT) and has finite sample guarantees for the test size under the null hypothesis unlike the LRT. We confirm theoretical properties in simulations that mimic the data set of our data application. We apply the method to a retinoblastoma data set and demonstrate a parent-of-origin effect. Weighted plausibility also has applications in high-dimensional data analysis and P-values for penalized regression models can be derived. We demonstrate superior performance as compared to a data-splitting procedure in a simulation study. We apply weighted plausibility to a high-dimensional gene expression, case-control prostate cancer data set. We discuss the flexibility of the approach by relating weighted plausibility to targeted learning, the bootstrap, and sparsity selection.

Citations