vix.ing · top · new · best · stats · spec

Empirical Null Estimation using Discrete Mixture Distributions and its\n Application to Protein Domain Data

2016/08/25 by Iris Ivy Gauran, Gauran, Iris Ivy, Junyong Park +14
Biochemistry, Genetics and Molecular Biology · #62F03 #62H12 #62H15 #FOS: Computer and information sciences #Gene expression and cancer classification #Genetic Associations and Epidemiology #Methodology (stat.ME) #Molecular Biology Techniques and Applications #RNA and protein synthesis mechanisms

paper · pdf · doi:10.48550/arxiv.1608.07204

openalex publication_date 2016/08/25 · openalex created_date 2022/10/02 · openalex updated_date 2026/07/28

Abstract

In recent mutation studies, analyses based on protein domain positions are\ngaining popularity over gene-centric approaches since the latter have\nlimitations in considering the functional context that the position of the\nmutation provides. This presents a large-scale simultaneous inference problem,\nwith hundreds of hypothesis tests to consider at the same time. This paper aims\nto select significant mutation counts while controlling a given level of Type I\nerror via False Discovery Rate (FDR) procedures. One main assumption is that\nthere exists a cut-off value such that smaller counts than this value are\ngenerated from the null distribution. We present several data-dependent methods\nto determine the cut-off value. We also consider a two-stage procedure based on\nscreening process so that the number of mutations exceeding a certain value\nshould be considered as significant mutations. Simulated and protein domain\ndata sets are used to illustrate this procedure in estimation of the empirical\nnull using a mixture of discrete distributions.\n

Citations

Related