Summary statistics knockoff inference empowers identification of putative causal variants in genome-wide association studies
He, Z.; Liu, L.; Belloy, M. E.; Le Guen, Y.; Sossin, A.; Liu, X.; Qi, X.; Ma, S.; Wyss-Coray, T.; Tang, H.; Sabatti, C.; Candes, E.; Greicius, M. D.; Ionita-Laza, I.
Show abstract
Recent advances in genome sequencing and imputation technologies provide an exciting opportunity to comprehensively study the contribution of genetic variants to complex phenotypes. However, our ability to translate genetic discoveries into mechanistic insights remains limited at this point. In this paper, we propose an efficient knockoff-based method, GhostKnockoff, for genome-wide association studies (GWAS) that leads to improved power and ability to prioritize putative causal variants relative to conventional GWAS approaches. The method requires only Z-scores from conventional GWAS and hence can be easily applied to enhance existing and future studies. The method can also be applied to meta-analysis of multiple GWAS allowing for arbitrary sample overlap. We demonstrate its performance using empirical simulations and two applications: (1) analysis of 1,403 binary phenotypes from the UK Biobank data in 408,961 samples of European ancestry, and (2) a meta-analysis for Alzheimers disease (AD) comprising nine overlapping large-scale GWAS, whole-exome and whole-genome sequencing studies. The UK Biobank analysis demonstrates superior performance of the proposed method compared to conventional GWAS in both statistical power (2.05-fold more discoveries) and localization of putative causal variants at each locus (46% less proxy variants due to linkage disequilibrium). The AD meta-analysis identified 55 risk loci (including 31 new loci) with ~70% of the proximal genes at these loci showing suggestive signal in downstream single-cell transcriptomic analyses. Our results demonstrate that GhostKnockoff can identify putatively functional variants with weaker statistical effects that are missed by conventional association tests.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A method to map and interpret pleiotropic loci using summary statistics of multiple traits 98%
- mBAT-combo: a more powerful test to detect gene-trait associations from GWAS data 97%
- Bayesian transcriptome-wide association study method leveraging both cis- and trans- eQTL information through summary statistics 97%
Similar papers in this journal
- Scalable generalized linear mixed model for region-based association tests in large biobanks and cohorts 97%
- Computationally efficient whole genome regression for quantitative and binary traits 97%
- A new method for multi-ancestry polygenic prediction improves performance across diverse populations 96%
Similar papers in this journal
- Identification of putative causal loci in whole-genome sequencing data via knockoff statistics 99%
- Testing and controlling for horizontal pleiotropy with the probabilistic Mendelian randomization in transcriptome-wide association studies 97%
- Quantifying portable genetic effects and improving cross-ancestry genetic prediction with GWAS summary statistics 97%
Similar papers in this journal
- Identity-by-descent mapping using multi-individual IBD with genome-wide multiple testing adjustment 96%
- RetroFun-RVS: a retrospective family-based framework for rare-variant analysis incorporating functional annotations 95%
- Meta-MultiSKAT: Multiple phenotype meta-analysis for region-based association test 95%
Similar papers in this journal
- Incorporating family disease history and controlling case-control imbalance for population based genetic association studies 96%
- Summary statistics from large-scale gene-environment interaction studies for re-analysis and meta-analysis 95%
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.