Back

The expected polygenic risk score (ePRS) framework: an equitable metric for quantifying polygenetic risk via modeling of ancestral makeup

Huang, Y.-J.; Kurniansyah, N.; Goodman, M. O.; Spitzer, B. W.; Wang, J.; Stilp, A. M.; Laurie, C.; de Vries, P. S.; Chen, H.; Min, Y.-I.; Sims, M.; Peloso, G. M.; Guo, X.; Bis, J. C.; Brody, J. A.; Raffield, L. M.; Smith, J. A.; Zhao, W.; Rotter, J. I.; Rich, S. S.; Redline, S.; Fornage, M.; Kaplan, R.; Franceschini, N.; Levy, D.; Morrison, A. C.; Boerwinkle, E.; Smith, N. L.; Kooperberg, C.; Psaty, B. M.; Zoellner, S.; Sofer, T.

2024-03-06 genetic and genomic medicine
10.1101/2024.03.05.24303738 medRxiv
Show abstract

Polygenic risk scores (PRSs) depend on genetic ancestry due to differences in allele frequencies between ancestral populations. This leads to implementation challenges in diverse populations. We propose a framework to calibrate PRS based on ancestral makeup. We define a metric called "expected PRS" (ePRS), the expected value of a PRS based on ones global or local admixture patterns. We further define the "residual PRS" (rPRS), measuring the deviation of the PRS from the ePRS. Simulation studies confirm that it suffices to adjust for ePRS to obtain nearly unbiased estimates of the PRS-outcome association without further adjusting for PCs. Using the TOPMed dataset, the estimated effect size of the rPRS adjusting for the ePRS is similar to the estimated effect of the PRS adjusting for genetic PCs. Similarly, we applied the ePRS framework to six cardiovascular-related traits in the All of Us dataset, and the results are consistent with those from the TOPMed analysis. The ePRS framework can protect from population stratification in association analysis and provide an equitable strategy to quantify genetic risk across diverse populations.

Published in Nature Communications (predicted rank #6) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.