Computationally efficient meta-analysis of gene-based tests using summary statistics in large-scale genetic studies
Joseph, T.; Mbatchou, J.; Ghosh, A.; Marcketta, A.; Gillies, C.; Tang, J.; Nakka, P.; Zhang, X.; Kosmicki, J.; Sidore, C.; Gurski, L.; Regeneron Genetics Center, ; Ghoussaini, M.; Ferreira, M. A. R.; Abecasis, G.; Marchini, J.
Show abstract
Meta-analysis of gene-based tests using single variant summary statistics is a powerful strategy for associating genes with disease. However, current approaches require sharing the covariance matrix between variants for each study and trait of interest. For large-scale studies with many phenotypes, these matrices can be cumbersome to calculate, store, and share. To address this challenge, we present REMETA, an efficient tool for meta-analysis of gene-based tests. REMETA uses a single sparse covariance reference file per study that is rescaled for each phenotype using single variant summary statistics. We develop methods to apply REMETA to binary traits with case-control imbalance, and estimate allele frequencies, genotype counts and effect sizes of burden tests. We demonstrate the performance and advantages of our approach via meta-analysis of 5 traits in 469,376 samples in UK Biobank. The open-source REMETA software tools and framework will facilitate meta-analysis across large scale exome sequencing studies from diverse studies that cannot easily be brought together.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Leveraging information between multiple population groups and traits improves fine-mapping resolution 98%
- Flashfm: A Flexible and Shared Information Fine-mapping Approach for Multiple Quantitative Traits 97%
- Identification of putative causal loci in whole-genome sequencing data via knockoff statistics 97%
Similar papers in this journal
Similar papers in this journal
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 97%
- Incorporating family disease history and controlling case-control imbalance for population based genetic association studies 96%
- Summary statistics from large-scale gene-environment interaction studies for re-analysis and meta-analysis 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.