Back

In silico identification of putative causal genetic variants

He, Z.; Chu, B.; Yang, J.; Gu, J.; Chen, Z.; Liu, L.; Morrison, T.; Belloy, M. E.; Qi, X.; Hejazi, N.; Mathur, M.; Le Guen, Y.; Tang, H.; Hastie, T.; Ionita-laza, I.; Sabatti, C.; Candes, E.

2024-03-03 genetics
10.1101/2024.02.28.582621 bioRxiv
Show abstract

Understanding the causal genetic architecture of complex phenotypes will fuel future research into disease mechanisms and potential therapies. Here, we illustrate the power of a novel framework: it detects, starting from summary statistics, and across the entire genome, sets of variants that carry non-redundant information on the phenotypes and are therefore more likely to be causal in a biological sense. The approach, implemented in open-source software, is also computationally efficient, requiring less than 15 minutes on a single CPU to perform genome-wide analysis. Through extensive genome-wide simulation studies, we show that the method can substantially outperform existing methods in false discovery rate control, statistical power and various fine-mapping criteria. In applications to a meta-analysis of ten large-scale genetic studies of Alzheimers disease (AD), we identified 82 loci associated with AD, including 37 additional loci missed by conventional GWAS pipeline. Massively parallel reporter assays and CRISPR-Cas9 experiments have confirmed the functionality of the putative causal variants our method points to. Finally, we retrospectively analyzed summary statistics from 67 large-scale GWAS for a variety of phenotypes. Results reveal the methods capacity to robustly discover additional loci for polygenic traits and pinpoint potential causal variants underpinning each locus beyond conventional GWAS pipeline, contributing to a deeper understanding of complex genetic architectures in post-GWAS analyses.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.