Large-scale composite hypothesis testing for omics analyses.
De Walsche, A.; Gauthier, F.; Charcosset, A.; Mary-Huard, T.
Show abstract
Composite hypothesis testing using summary statistics is a well-established approach for assessing the effect of a single marker or gene across multiple traits or omics levels. Numerous procedures have been developed for this task and have been successfully applied to identify complex patterns of association between traits, conditions, or phenotypes. However, existing methods often struggle with scalability in large datasets or fail to account for dependencies between traits or omics levels, limiting their ability to control false positives effectively. To overcome these challenges, we present the qch_copula approach, which integrates mixture models with a copula function to capture dependencies between traits or omics, and provides rigorously defined p-values for any composite hypothesis. Through a comprehensive benchmark against eight state-of-the-art methods, we demonstrate that qch_copula controls Type I error rates effectively while enhancing the detection of joint association patterns. Compared to other mixture model-based approaches, our method notably reduces memory usage during the EM algorithm, allowing the analysis of up to 20 traits and 105 - 106 markers. The effectiveness of qch_copula is further validated through two application cases in human and plant genetics. The method is available in the R package qch, accessible on CRAN.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An exact, unifying framework for region-based association testing in family-based designs, including higher criticism approaches, SKATs, multivariate and burden tests 95%
- networkGWAS:A network-based approach to discover genetic associations 95%
- Subset scanning for multi-trait analysis using GWAS summary statistics 95%
Similar papers in this journal
Similar papers in this journal
- Taking population stratification into account by local permutations in rare-variant association studies on small samples 95%
- A robust association test leveraging unknown genetic interactions: Application to cystic brosis lung disease 95%
- Quantifying posterior effect size distribution of susceptibility loci by common summary statistics 94%
Similar papers in this journal
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 95%
- BayesKAT: Bayesian Optimal Kernel-based Test for genetic association studies reveals joint genetic effects in complex diseases 95%
- Efficient test for deviation from Hardy Weinberg Equilibrium with known or ambiguous typing in highly polymorphic loci 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.