2018/01/11 by Martin Jinye Zhang, Meisam Razaviyayn, Zhang, Martin J. +3
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Bayesian Methods and Mixture Models #FOS: Mathematics #Gene expression and cancer classification #Statistical Methods and Inference #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.1801.04005
openalex publication_date 2018/01/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Comparing two groups under different conditions is ubiquitous in the biomedical sciences. In many cases, samples from the two groups can be naturally paired; for example a pair of samples may come from the same individual under the two conditions. However samples across different individuals may be highly heterogeneous. Traditional methods often ignore such heterogeneity by assuming the samples are identically distributed. In this work, we study the problem of comparing paired heterogeneous data by modeling the data as Gaussian distributed with different parameters across the samples. We show that in the minimax setting where we want to maximize the worst-case power, the sign test, which only uses the signs of the differences between the paired sample, is optimal in the one-sided case and near optimal in the two-sided case. The superiority of the sign test over other popular tests for paired heterogeneous data is demonstrated using both synthetic data and a real-world RNA-Seq dataset.