Back

A Principled Framework for Using Correlated Traits to Improve Risk Prediction

Akey, J. M.; Bierman, R.; Zhang, K.

2026-08-27 genomics
10.64898/2026.08.23.746504 bioRxiv
Show abstract

Although many complex phenotypes and diseases are influenced by shared genetic and environmental factors, risk prediction methods typically rely on genetic information from a single trait, leaving a rich source of predictive information largely unexploited. Phenotypic correlations can potentially be used to improve the accuracy of polygenic scores (PGS), but the conditions under which correlated traits meaningfully enhance prediction remain poorly understood. Here, we develop a general theoretical and simulation framework that quantifies the extent to which correlated "helper" traits improve predictive accuracy and identifies the factors that determine the magnitude of these gains. We show that helper traits can substantially improve predictive accuracy, with the magnitude of these gains governed by baseline model performance, genetic and environmental correlations, and the heritability of the target and helper traits, providing principled guidance for helper-trait selection. Paradoxically, when the target trait is itself weakly heritable, helper traits need not be highly heritable to substantially improve the accuracy of PGS, because low-heritability traits can still capture non-redundant environmental factors shared with the target trait. We empirically evaluated the use of helper traits by developing PGS models to predict type 2 diabetes using data from the UK Biobank. Helper traits substantially improved predictive accuracy relative to a single-trait PGS (AUC-ROC = 0.907 versus 0.677) and achieved performance comparable to models that use HbA1c (AUC-ROC = 0.889), the current clinical gold-standard biomarker. Our results establish a general theoretical and practical framework for exploiting correlated traits to improve polygenic prediction, provide principled guidance for selecting informative helper traits, and demonstrate how shared genetic and environmental architecture can be leveraged to substantially increase predictive accuracy. Furthermore, we developed an interactive web application to estimate the expected gain in accuracy from candidate helper traits using empirically measurable quantities.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.