Back

Ancestry Calibration of Polygenic Risk Scores Improves Risk Stratification and Effect Estimation in African American Adults

Vargas, L. B.; Meyer, M. C.; Konigsberg, I. R.; Kakar, A.; Carry, P. M.; Polygenic Risk Methods in Diverse Populations (PRIMED) Consortium Methods Working Group, ; Li, Y.; Jones, A. C.; Tiwari, H. K.; Srinivasasainagendra, V.; Armstrong, N. D.; Kenny, E. E.; Pasaniuc, B.; Irvin, M. R.; Cho, M. H.; Stanislawski, M. A.; Raghavan, S.; Shortt, J. A.; Lange, L. A.; Lange, E. M.

2025-06-18 genetic and genomic medicine
10.1101/2025.06.18.25329573 medRxiv
Show abstract

Polygenic risk score (PRS) distributions vary across populations, complicating PRS risk assessment. We evaluated the impact of post-hoc PRS calibration according to individualized genetic ancestry estimates on PRS performance using two large multi-ethnic PRS for type 2 diabetes (T2D) (PRST2D) and height (PRSheight), in 8,841 African American (AA) individuals from the Reasons for Geographic and Racial Differences in Stroke (REGARDS) study. We calibrated each participants score as a function of estimated genetic similarity to the Yoruba (GSYRI) cohort in the 1000 Genomes Project. Uncalibrated PRSs were significantly skewed by GSYRI. After calibration, 33.6% of individuals in the top decile of PRST2D were reclassified and performance in the top PRST2D decile improved from an OR of 7.97 [6.31-10.13] to 10.77 [8.41- 13.91] when compared to the lowest decile. Similarly, 55.0% of individuals in the top PRSheight decile were reclassified with GSYRI calibration. The calibrated PRSheight showed higher correlation with height (from 0.24 to 0.32, p<10-7), and increased mean height in the top PRSheight decile (p=5.7x10-5) when compared to the uncalibrated PRSheight. Lastly, we show that evaluating uncalibrated PRS while adjusting for GSYRI in regression models can lead to inflated and unstable effect size estimates for both the PRS and GSYRI.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.