2019/12/12 by Paramita Chakraborty, Chong Ma, Chakraborty, Paramita +6
Biochemistry, Genetics and Molecular Biology · Mathematics · #Applications (stat.AP) #Computation (stat.CO) #FOS: Computer and information sciences #Gene expression and cancer classification #Molecular Biology Techniques and Applications #Statistical Methods in Clinical Trials
paper · pdf · doi:10.48550/arxiv.1912.06030
openalex publication_date 2019/12/12 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
In large scale multiple testing problems, a two-class empirical Bayes\napproach can be used to control the false discovery rate (Fdr) for the entire\narray of hypotheses under study. A sample splitting step is incorporated to\nmodify that approach where one part of the data is used for model fitting and\nthe other part for detecting the significant cases by a screening technique\nfeaturing the empirical Bayes mode of Fdr control. Cases with high detection\nfrequency across repeated random sample splits are considered true discoveries.\nA critical detection frequency is set to control the overall false discovery\nrate. The proposed method helps to balance out unwanted sources of variation\nand addresses potential statistical overfitting of the core empirical model by\ncross-validation through resampling. Further, concurrent detection frequencies\nare used to provide visual tools to explore the inter-relationship between\nsignificant cases. The methodology is illustrated using a microarray data set,\nRNA-sequencing data set, and several simulation studies. A power analysis is\npresented to understand the efficiency of the proposed method.\n