Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GxE PRS: Genotype-environment interaction in polygenic risk score models for quantitative and binary traits 93%
- Statistical learning for sparser fine-mapped polygenic models: the prediction of LDL-cholesterol 93%
- Statistics to prioritize rare variants in family-based sequencing studies with disease subtypes 92%
Similar papers in this journal
- A parametric bootstrap approach for computing confidence intervals for genetic correlations with application to genetically-determined protein-protein networks 91%
- A simple approach for multiple observations improves power to detect genetic effects and genomic prediction accuracy. 91%
- Leveraging Global Genetics Resources to Enhance Polygenic Prediction Across Ancestrally Diverse Populations 90%
Similar papers in this journal
- A two-step approach to testing overall effect of gene-environment interaction for multiple phenotypes 93%
- An exact, unifying framework for region-based association testing in family-based designs, including higher criticism approaches, SKATs, multivariate and burden tests 92%
- Subset scanning for multi-trait analysis using GWAS summary statistics 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.