2023/05/09 by Zhanrui Cai, Cai, Zhanrui
Computer Science · Mathematics · #Bayesian Methods and Mixture Models #FOS: Biological sciences #FOS: Computer and information sciences #Genomics (q-bio.GN) #Methodology (stat.ME) #Statistical Methods and Bayesian Inference #Statistical Methods and Inference
paper · pdf · doi:10.48550/arxiv.2305.05714
openalex publication_date 2023/05/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Model selection is critical in the modern statistics and machine learning community. However, most existing works do not apply to heavy-tailed data, which are commonly encountered in real applications, such as the single-cell multiomics data. In this paper, we propose a rank-sum based approach that outputs a confidence set containing the optimal model with guaranteed probability. Motivated by conformal inference, we developed a general method that is applicable without moment or tail assumptions on the data. We demonstrate the advantage of the proposed method through extensive simulation and a real application on the COVID-19 genomics dataset (Stephenson et al., 2021). To perform the inference on rank-sum statistics, we derive a general Gaussian approximation theory for high dimensional two-sample U-statistics, which may be of independent interest to the statistics and machine learning community.