Boosting the power of rare variant association studies by imputation using large-scale sequencing population
Dai, J.; Zhang, Y.; Li, Z.; Li, H.; Du, S.; You, D.; Zhang, R.; Zhao, Y.; Liu, Z.; Christiani, D. C.; Chen, F.; Shen, S.
Show abstract
Rare variants can explain part of the heritability of complex traits that are ignored by conventional GWASs. The emergence of large-scale population sequencing data provides opportunities to study rare variants. However, few studies systematically evaluate the extent to which imputation using sequencing data can improve the power of rare variant association studies. Using whole genome sequencing (WGS) data (n = 150,119) as the ground truth, we described the landscape and evaluated the consistency of rare variants in SNP array (n = 488,377) imputed from TOPMed or HRC+UK10K in the UK Biobank, respectively. The TOPMed imputation covered more rare variants, and its imputation quality could reach 0.5 for even extremely rare variants. TOPMed-imputed data was closer to WGS in all MAC intervals for three ethnicities (average Cramers V>0.75). Furthermore, association tests were performed on 30 quantitative and 15 binary traits. Compared to WGS data, the identified rare variants in TOPMed-imputed data increased 27.71% for quantitative traits, while it could be improved by [~]10-fold for binary traits. In gene-based analysis, the signals in TOPMed-imputed data increased 111.45% for quantitative traits, and it identified 15 genes in total, while WGS only found 6 genes for binary traits. Finally, we harmonized SNP array and WGS data for lung cancer and epithelial ovarian cancer. More variants and genes could be identified than from WGS data alone, such as BRCA1, BRCA2, and CHRNA5. Our findings highlighted that incorporating rare variants imputed from large-scale sequencing populations could greatly boost the power of GWAS.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Single cell sequencing analysis uncovers genetics-influenced CD16+monocytes and memory CD8+T cells involved in severe COVID-19 94%
- Illuminating links between cis-regulators and trans-acting variants in the human prefrontal cortex 93%
- Whole-genome reference panel of 1,781 Northeast Asians improves imputation accuracy of rare and low-frequency variants 93%
Similar papers in this journal
- Utilizing Non-Invasive Prenatal Test Sequencing Data Resource for Human Genetic Investigation 95%
- Taiwan Biobank: a rich biomedical research database of the Taiwanese population 95%
- Genome-wide study on 72,298 Korean individuals in Korean biobank data for 76 traits identifies hundreds of novel loci 95%
Similar papers in this journal
- SUMMIT: An integrative approach for better transcriptomic data imputation improves causal gene identification 95%
- Projecting genetic associations through gene expression patterns highlights disease etiology and drug mechanisms 94%
- Accounting for genetic effect heterogeneity in fine-mapping and improving power to detect gene-environment interactions with SharePro 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.