Back

Large-scale evaluation of proteomic and polygenic risk scores reveals complementary contributions to incident disease prediction

Woerner, J.; Westbrook, T. M.; Joo, J.; Shivakumar, M.; Venkatesh, R.; Cherlin, T.; JUNG, S.; Jeong, S.; Maseda, D.; McKeague, M.; Shwetank, F.; Ionita, M.; Wagenaar, J.; Abramowitz, S. A.; Verma, A.; Zhao, B.; Lee, S.; Damrauer, S. M.; Levin, M. M.; Heo, S.-J.; Cappola, T.; Rader, D. J.; Day, S. M.; Deo, R.; Gelfand, J.; Ramessur, R.; Guerraty, M.; Verma, S. S.; Pasaniuc, B.; Ritchie, M. D.; Apostolidis, S. A.; Greenplate, A.; Wherry, E. J. A.; Penn Medicine Biobank, ; Nam, Y.; Kim, D.

2025-07-11 genetic and genomic medicine
10.1101/2025.07.10.25331242 medRxiv
Show abstract

Plasma proteins capture dynamic physiological processes and may offer more immediate insight into disease risk than static genetic predictors. We evaluated the predictive utility of proteomic risk scores (ProRS) versus polygenic risk scores (PRS) across 301 phenotypes in 39,843 participants from the UK Biobank Pharma Proteomics Project. ProRS, trained on prevalent cases, were tested for incident disease and benchmarked against PRS derived from genome-wide association statistics. Among 268 phenotypes with informative signals, ProRS outperformed PRS in 88% of traits (a median C-index improvement of 9.6%), showing strongest gains for circulatory, metabolic, and immune conditions. Combined models further improved prediction, particularly for traits with higher heritability. Longitudinal analyses showed that ProRS values were elevated years before diagnosis. External validation in 841 Penn Medicine BioBank participants confirmed consistent performance and transferability, with AUC improvements up to 4.18% over PRS alone. Plasma proteomic profiling provides complementary, temporally responsive information that enhances individual-level disease prediction.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.