2024/04/08 by Pengfei Lyu, Xianyang Zhang, Lyu, Pengfei +3 · 1 voice
Business, Management and Accounting · Computer Science · Decision Sciences · #Big Data and Business Intelligence #Data Quality and Management #Software Reliability and Analysis Research #stat.ME
paper · pdf · doi:10.48550/arxiv.2404.05808
openalex publication_date 2024/04/08 · openalex created_date 2024/04/11 · openalex updated_date 2026/07/28
Testing composite null hypotheses is fundamental to many scientific applications, including mediation and replicability analyses, and becomes particularly challenging in high-throughput settings involving tens of thousands of features. Existing high-dimensional composite null hypotheses testing often ignores the dependence structure among features, leading to overly conservative or liberal results. To address this limitation, we develop a four-state hidden Markov model (HMM) for bivariate p-value sequences arising from two-study replicability analysis. This model captures local dependence among features and accommodates study-specific heterogeneity. Based on the HMM, we propose a multiple testing procedure that asymptotically controls the false discovery rate (FDR). Extending this framework to more than two studies is computationally intensive, with complexity growing exponentially in the number of studies n. To address this scalability issue, we introduce a novel e-value framework that reduces computational complexity to quadratic in n, while preserving asymptotic FDR control. Extensive simulations demonstrate that our method achieves higher power than existing approaches at the same FDR levels. When applied to genome-wide association studies (GWAS), the proposed approach identifies replicable SNP-level signals that are not detected at the same significance threshold by competing methods.