Imputation and polygenic score performances of human genotyping arrays in diverse populations
Nguyen, D. T.; Tran, T.; Tran, M.; Tran, K.; Pham, D.; Duong, N. T.; Nguyen, Q.; Vo, N. S.
Show abstract
Regardless of the overwhelming use of next-generation sequencing technologies, microarray-based genotyping combined with the imputation of untyped variants remains a cost-effective means to interrogate genetic variations across the human genome. This technology is widely used in genome-wide association studies (GWAS) at bio-bank scales, and more recently, in polygenic score (PGS) analysis to predict and to stratify disease risk. Over the last decade, human genotyping arrays have undergone a tremendous growth in both number, and content making a comprehensive evaluation of their performances became more important. Here, we performed a comprehensive performance assessment for 23 available human genotyping arrays in 6 ancestry groups using diverse public, and in-house datasets. The analyses focus on performance estimation of derived imputation (in terms of accuracy and coverage) and PGS (in term of concordance to PGS estimated from whole genome sequencing data) in three different traits and diseases. We found that the arrays with a higher number of SNPs are not necessarily the ones with higher imputation performance, but the arrays that are well-optimized for the targeted population could provide very good imputation performance. In addition, PGS estimated by imputed SNP array data is highly correlated to PGS estimated by whole genome sequencing data in most of cases. When optimal arrays are used, the correlations of key PGS metrics between two types of data can be higher than 0.97, but interestingly, arrays with high density can result in lower PGS performance. Our results suggest the importance of properly selecting a suitable genotyping array for PGS applications. Finally, we developed a web tool that provide interactive analyses of tag SNP contents and imputation performance based on population and genomic regions of interest. This study would act as a practical guide for researchers to design their genotyping arrays-based studies. The tool is available at: https://genome.vinbigdata.org/tools/saa/
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GMQN: A reference-based method for correcting batch effects as well as probes bias in HumanMethylation BeadChip 93%
- Hardy-Weinberg Equilibrium in the Large Scale Genomic Sequencing Era 93%
- Localization of balanced chromosome translocation breakpoints by long-read sequencing on the Oxford Nanopore platform 92%
Similar papers in this journal
- Similarity and diversity of genetic architecture for complex traits between East Asian and European populations 93%
- A Nextflow pipeline for molecular quantitative trait loci mapping in small sample size datasets with an application in Atlantic salmon 93%
- An individualized Bayesian method for estimating genomic variants of hypertension 93%
Similar papers in this journal
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 94%
- LDBlockShow: a fast and convenient tool for visualizing linkage disequilibrium and haplotype blocks based on variant call format files 93%
- A novel haplotype-based eQTL approach identifies genetic associations not detected through conventional SNP-based methods 93%
Similar papers in this journal
- Principal component analysis- and tensor decomposition-based unsupervised feature extraction to select more reasonable differentially methylated cytosines: Optimization of standard deviation versus state-of-the-art methods 92%
- A genome-wide epistatic network underlies the molecular architecture of continuous color variation of body extremities: a rabbit model 91%
- Development and validation of a combined species SNP array for the European seabass (Dicentrarchus labrax) and gilthead seabream (Sparus aurata) 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.