A Principled Framework for Using Correlated Traits to Improve Risk Prediction
Akey, J. M.; Bierman, R.; Zhang, K.
Show abstract
Although many complex phenotypes and diseases are influenced by shared genetic and environmental factors, risk prediction methods typically rely on genetic information from a single trait, leaving a rich source of predictive information largely unexploited. Phenotypic correlations can potentially be used to improve the accuracy of polygenic scores (PGS), but the conditions under which correlated traits meaningfully enhance prediction remain poorly understood. Here, we develop a general theoretical and simulation framework that quantifies the extent to which correlated "helper" traits improve predictive accuracy and identifies the factors that determine the magnitude of these gains. We show that helper traits can substantially improve predictive accuracy, with the magnitude of these gains governed by baseline model performance, genetic and environmental correlations, and the heritability of the target and helper traits, providing principled guidance for helper-trait selection. Paradoxically, when the target trait is itself weakly heritable, helper traits need not be highly heritable to substantially improve the accuracy of PGS, because low-heritability traits can still capture non-redundant environmental factors shared with the target trait. We empirically evaluated the use of helper traits by developing PGS models to predict type 2 diabetes using data from the UK Biobank. Helper traits substantially improved predictive accuracy relative to a single-trait PGS (AUC-ROC = 0.907 versus 0.677) and achieved performance comparable to models that use HbA1c (AUC-ROC = 0.889), the current clinical gold-standard biomarker. Our results establish a general theoretical and practical framework for exploiting correlated traits to improve polygenic prediction, provide principled guidance for selecting informative helper traits, and demonstrate how shared genetic and environmental architecture can be leveraged to substantially increase predictive accuracy. Furthermore, we developed an interactive web application to estimate the expected gain in accuracy from candidate helper traits using empirically measurable quantities.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning Gene Networks Underlying Clinical Phenotypes Using SNP Perturbations 92%
- Efficient and Flexible Integration of Variant Characteristics in Rare Variant Association Studies Using Integrated Nested Laplace Approximation 92%
- Enabling interpretable machine learning for biological data with reliability scores 91%
Similar papers in this journal
- Co-expression-wide association studies link genetically regulated interactions with complex traits 93%
- Probabilistic inference of the genetic architecture underlying functional enrichment of complex traits 93%
- The expected polygenic risk score (ePRS) framework: an equitable metric for quantifying polygenetic risk via modeling of ancestral makeup 92%
Similar papers in this journal
- Polygenic Health Index, General Health, Pleiotropy, Embryo Selection and Disease Risk 94%
- Optimization of Multi-Ancestry Polygenic Risk Score Disease Prediction Models 93%
- Sibling Variation in Phenotype and Genotype: Polygenic Trait Distributions and DNA Recombination Mapping with UK Biobank and IVF Family Data 92%
Similar papers in this journal
- The Causal Pivot: A Structural Approach to Genetic Heterogeneity and Variant Discovery in Complex Diseases 94%
- Estimating disease heritability from complex pedigrees allowing for ascertainment and covariates 93%
- Welch-weighted Egger regression reduces false positives due to correlated pleiotropy in Mendelian randomization 93%
Similar papers in this journal
- Characterization of direct and/or indirect genetic associations for multiple traits in longitudinal studies of disease progression 93%
- Hierarchical clustering of gene-level association statistics reveals shared and differential genetic architecture among traits in the UK Biobank 92%
- A Bayesian Approach to Correcting the Attenuation Bias of Regression Using Polygenic Risk Score 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.