Imputation and polygenic score performance of low coverage whole-genome sequencing and genotyping arrays in diverse human populations
Nguyen, P. T.; Nguyen, V. T.; Nguyen, D. T.; Duong, H. H.-T.
Show abstract
Genome-wide association studies and polygenic score analysis rely on large-scale genotypic data, traditionally obtained through SNP arrays and imputation. However, low coverage whole-genome sequencing has emerged as a promising alternative. This study presents a comprehensive comparison of imputation accuracy and polygenic score performance between eight high-performance genotyping arrays and six low coverage whole-genome sequencing coverage levels (0.5-2x) across diverse populations. We analyze data from 2,504 individuals in the 1000 Genomes Project using a 10-fold cross-imputation strategy to evaluate imputation accuracy and polygenic score performance for four complex traits. Our results demonstrate that low-pass whole-genome sequencing performs competitively with population-specific arrays in both imputation accuracy and polygenic score estimation. Interestingly, low coverage whole-genome sequencing shows superior performances compared to arrays in underrepresented populations and for rare and low-frequency variants. Our findings suggest that low coverage whole-genome sequencing offers a flexible and powerful alternative to genotyping arrays for large-scale genetic studies, particularly in diverse or underrepresented populations.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Nextflow pipeline for molecular quantitative trait loci mapping in small sample size datasets with an application in Atlantic salmon 94%
- Efficient blockLASSO for Polygenic Scores with Applications to All of Us and UK Biobank 94%
- Similarity and diversity of genetic architecture for complex traits between East Asian and European populations 93%
Similar papers in this journal
Similar papers in this journal
- A reference panel for linkage disequilibrium and genotype imputation using whole-genome sequencing data from 2,680 participants across India 94%
- Inverted genomic regions between reference genome builds in humans impact imputation accuracy and decrease the power of association testing 94%
- Leveraging Global Genetics Resources to Enhance Polygenic Prediction Across Ancestrally Diverse Populations 93%
Similar papers in this journal
- Increasing calling accuracy, coverage, and read depth in sequence data by the use of haplotype blocks 94%
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 94%
- Improving polygenic prediction from summary data by learning patterns of effect sharing across multiple phenotypes. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.