vix.ing · top · new · best · stats · spec

Variable Selection for Multi-Source Count Data with Controlled False Discovery Rate

2024/11/28 by Shanjun Mao, Tang, Shan, Stephens Ma +4
Computer Science · Medicine · #Applications (stat.AP) #Bayesian Modeling and Causal Inference #Data Stream Mining Techniques #Data-Driven Disease Surveillance #FOS: Computer and information sciences #Methodology (stat.ME)

paper · pdf · doi:10.48550/arxiv.2411.18986

openalex publication_date 2024/11/28 · openalex created_date 2024/12/05 · openalex updated_date 2026/07/28

Abstract

The rapid generation of complex, highly skewed, and zero-inflated multi-source count data poses significant challenges for variable selection, particularly in biomedical domains like tumor development and metabolic dysregulation. To address this, we propose a new variable selection method, Zero-Inflated Poisson-Gamma Simultaneous Knockoff (ZIPG-SK), specifically designed for multi-source count data. Our method leverages a gaussian copula based on the Zero-Inflated Poisson-Gamma (ZIPG) distribution to construct knockoffs that properly account for the properties of count data, including high skewness and zero inflation, while effectively incorporating covariate information. This framework enables the detection of common features across multi-source datasets with guaranteed false discovery rate (FDR) control. Furthermore, we enhance the power of the method by incorporating e-value aggregation, which effectively mitigates the inherent randomness in knockoff generation. Through extensive simulations, we demonstrate that ZIPG-SK significantly outperforms existing methods, achieving superior power across various scenarios. We validate the utility of our method on real-world colorectal cancer (CRC) and type 2 diabetes (T2D) datasets, identifying key variables whose characteristics align with established findings and simultaneously provide new mechanistic insights.

Related