2025/07/04 by Annaïg De Walsche, Franck Gauthier, Nathalie Boissot +2 · 1 voice
Biochemistry, Genetics and Molecular Biology · #Gene expression and cancer classification #Genetic Associations and Epidemiology #Genetic Mapping and Diversity in Plants and Animals
paper · doi:10.1093/nargab/lqaf118
openalex publication_date 2025/07/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract Composite hypothesis testing using summary statistics is a well-established approach for assessing the effect of a single marker or gene across multiple traits or omics levels. Numerous procedures have been developed for this task and have been successfully applied to identify complex patterns of association between traits, conditions, or phenotypes. However, existing methods often struggle with scalability in large datasets or fail to account for dependencies between traits or omics levels, limiting their ability to control false positives effectively. To overcome these challenges, we present the qchcopula approach, which integrates mixture models with a copula function to capture dependencies between traits or omics and provides rigorously defined P-values for any composite hypothesis. Through a comprehensive benchmark against eight state-of-the-art methods, we demonstrate that qchcopula controls Type I error rates effectively while enhancing the detection of joint association patterns. Compared to other mixture model-based approaches, our method notably reduces memory usage during the EM algorithm, allowing the analysis of up to 20 traits and 105−106 markers. The effectiveness of qchcopula is further validated through two application cases in human and plant genetics. The method is available in the R package qch, accessible on CRAN.