Polygenic scoring accuracy varies across the genetic ancestry continuum in all human populations
Ding, Y.; Hou, K.; Xu, Z.; Pimplaskar, A.; Petter, E.; Boulier, K.; Prive, F.; Vilhjalmsson, B. J.; Loohuis, L. O.; Pasaniuc, B.
Show abstract
Polygenic scores (PGS) have limited portability across different groupings of individuals (e.g., by genetic ancestries and/or social determinants of health), preventing their equitable use. PGS portability has typically been assessed using a single aggregate population-level statistic (e.g., R2), ignoring inter-individual variation within the population. Here we evaluate PGS accuracy at individual-level resolution, independent of its annotated genetic ancestries. We show that PGS accuracy varies between individuals across the genetic ancestry continuum in all ancestries, even within traditionally "homogeneous" genetic ancestry clusters. Using a large and diverse Los Angeles biobank (ATLAS, N= 36,778) along with the UK Biobank (UKBB, N= 487,409), we show that PGS accuracy decreases along a continuum of genetic ancestries in all considered populations and the trend is well-captured by a continuous measure of genetic distance (GD) from the PGS training data; Pearson correlation of -0.95 between GD and PGS accuracy averaged across 84 traits. When applying PGS models trained in UKBB "white British" individuals to European-ancestry individuals of ATLAS, individuals in the highest GD decile have 14% lower accuracy relative to the lowest decile; notably the lowest GD decile of Hispanic/Latino American ancestry individuals showed similar PGS performance as the highest GD decile of European ancestry ATLAS individuals. GD is significantly correlated with PGS estimates themselves for 82 out of 84 traits, further emphasizing the importance of incorporating the continuum of genetic ancestry in PGS interpretation. Our results highlight the need for moving away from discrete genetic ancestry clusters towards the continuum of genetic ancestries when considering PGS and their applications.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CADET: Enhanced transcriptome-wide association analyses in admixed samples using eQTL summary data 98%
- Estimating heritability explained by local ancestry and evaluating stratification bias in admixture mapping from summary statistics 97%
- Estimation of non-additive genetic variance in human complex traits from a large sample of unrelated individuals 97%
Similar papers in this journal
- Sparse haplotype-based fine-scale local ancestry inference at scale reveals recent selection on immune responses 97%
- Theoretical and empirical quantification of the accuracy of polygenic scores in ancestry divergent populations 97%
- Quantifying portable genetic effects and improving cross-ancestry genetic prediction with GWAS summary statistics 97%
Similar papers in this journal
- Calibrated prediction intervals for polygenic scores across diverse contexts 98%
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 97%
- Extremely sparse models of linkage disequilibrium in ancestrally diverse association studies 97%
Similar papers in this journal
- Evaluating genetic-ancestry inference from single-cell transcriptomic datasets 97%
- Multivariate adaptive shrinkage improves cross-population transcriptome prediction for transcriptome-wide association studies in underrepresented populations 96%
- Powerful eQTL mapping through low coverage RNA sequencing 95%
Similar papers in this journal
- MUSSEL: Enhanced Bayesian Polygenic Risk Prediction Leveraging Information across Multiple Ancestry Groups 96%
- Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases 96%
- Polymorphic short tandem repeats make widespread contributions to blood and serum traits 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.