Efficient identification of trait-associated loss-of-function variants in the UK Biobank cohort by exome-sequencing based genotype imputation
Zhang, L.; Yan, S.-S.; Ni, J.-J.; Pei, Y.-F.
Show abstract
The large-scale open access whole-exome sequencing (WES) data of the UK Biobank ~200,000 participants is accelerating a new wave of genetic association studies aiming to identify rare and functional loss-of-function (LoF) variants associated with a broad range of complex traits and diseases, however the community is in short of stringent replication of new associations. In this study, we proposed to merge the WES genotypes and the genome-wide genotyping (GWAS) genotypes of 167,000 UKB Caucasian participants into a combined reference panel, and then to impute 241,911 UKB Caucasian participants who had the GWAS genotypes only. We then proposed to use the imputed data to replicate association identified in the discovery WES sample. Using a leave-100-out imputation strategy in the reference panel, we showed that average imputation accuracy measure r2 is modest to high at LoF variants of all minor allele frequency (MAF) intervals including ultra-rare ones: 0.942 at MAF interval [1%, 50%], 0.807 at [0.1%, 1.0%), 0.805 at [0.01%, 0.1%), 0.664 at [0.001%, 0.01%) and 0.410 at (0, 0.001%). As applications, we studied single variant level and gene level associations of LoF variants with estimated heel BMD (eBMD) and 4 lipid traits: high-density-lipoprotein cholesterol (HDL-C), low-density-lipoprotein cholesterol (LDL-C), triglycerides (TG) and total cholesterol (TC). In addition to replicating dozens of previously reported genes such as MEPE for eBMD and PCSK9 for more than one lipid trait, the results also identified 2 novel gene-level associations: PLIN1 (cumulative MAF=0.10%, discovery BETA=0.38, P=1.20x10-13; replication BETA=0.25, P=1.03x10-6) and ANGPTL3 (cumulative MAF=0.10%, discovery BETA=-0.36, P=4.70x10-11; replication BETA=-0.30, P=6.60x10-11) for HDL-C, as well as one novel single variant level association (11:14843853:C:T, MAF=0.11%, discovery BETA=-0.31, P=2.70x10-9; replication BETA=-0.31, P=8.80x10-14, PDE3B) for TG. Our results highlighted the strength of WES based genotype imputation as well as provided useful imputed data within the UKB cohort.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Scalable Framework for Identifying Allelic Series from Summary Statistics 97%
- An allelic series rare variant association test for candidate gene discovery 96%
- Cancer PRSweb - an Online Repository with Polygenic Risk Scores (PRS) for Major Cancer Traits and Their Phenome-wide Exploration in Two Independent Biobanks 95%
Similar papers in this journal
- Genomic dissection of 43 serum urate-associated loci provides multiple insights into molecular mechanisms of urate control. 94%
- A data harmonization pipeline to leverage external controls and boost power in GWAS 94%
- Imputed Gene Expression Risk Scores: A Functionally Informed Component of Polygenic Risk 94%
Similar papers in this journal
- Common, low-frequency, rare, and ultra-rare coding variants contribute to COVID-19 severity 93%
- Multi-omics highlights ABO plasma protein as a causal risk factor for COVID-19 92%
- Whole genome sequencing of orofacial cleft trios from the Gabriella Miller Kids First Pediatric Research Consortium identifies a new locus on chromosome 21 91%
Similar papers in this journal
- Inverted genomic regions between reference genome builds in humans impact imputation accuracy and decrease the power of association testing 96%
- Evaluation of imputation performance of multiple reference panels in a Pakistani population 95%
- Pitfalls in performing genome-wide association studies on ratio traits 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.