Empowering GWAS Discovery through Enhanced Genotype Imputation
De Marino, A.; Mahmoud, A. A.; Bohn, S.; Lerga-Jaso, J.; Novkovic, B.; Manson, C.; Loguercio, S.; Terpolovsky, A.; Matushyn, M.; Torkamani, A.; Yazdi, P. G.
Show abstract
MotivationGenotype imputation is a powerful tool for inferring missing genotype data in large-scale genomic studies. Over the last two decades, multiple research groups have developed a number of imputation algorithms, which continue improving in speed and overall accuracy. However, accurate imputation of rare and infrequent variants remains a challenge. ResultsHere we present Selphi, a novel genotype imputation algorithm based on the Positional Burrow Wheeler Transform (PBWT) and a new heuristic method for haplotype selection based on identity by descent (IBD). When compared to state-of-the-art methods Beagle5.4, IMPUTE5, and Minimac4, Selphi showed a higher accuracy in 1000 Genome Project and TOPmed datasets, across all super-populations and allele frequencies. Similarly, Selphi performed better than Beagle5.4 in the UK Biobank dataset, which translated into improved GWAS discovery and more accurate polygenic risk scores. Selphis improvements in imputation accuracy, especially for rare and low frequency variants, promises to boost the power and accuracy of downstream genomic applications. Availability and implementationSelphi code is available at GitHub: https://github.com/selfdecode/rd-imputation-selphi. Additionally, we offer an applet that allows convenient testing of the Selphi code on the UKB RAP platform.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Meta-analysis fine-mapping is often miscalibrated at single-variant resolution 96%
- Polymorphic short tandem repeats make widespread contributions to blood and serum traits 96%
- Analysis across Taiwan Biobank, Biobank Japan and UK Biobank identifies hundreds of novel loci for 36 quantitative traits 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.