One score to rule them all: regularized ensemble polygenic risk prediction with GWAS summary statistics
Zhao, Z.; Dorn, S.; Wu, Y.; Yang, X.; Jin, J.; Lu, Q.
Show abstract
Ensemble learning has been increasingly popular for boosting the predictive power of polygenic risk scores (PRS), with almost every recent multi-ancestry PRS approach employing ensemble learning as a final step. Existing ensemble approaches rely on individual-level data for model training, which severely limits their real-world applications, especially in non-European populations without sufficient genomic samples. Here, we introduce a statistical framework to construct regularized ensemble PRS, which allows us to combine a large number of candidate PRS models using only summary statistics from genome-wide association studies. We demonstrate its robust and substantial improvement over many existing PRS models in both within- and cross-ancestry applications. We believe this is truly "one score to rule them all" due to its capability to continuously combine newly developed PRS models with existing models to improve prediction performance, which makes it a universal approach that should always be employed in future PRS applications.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 99%
- A new method for multi-ancestry polygenic prediction improves performance across diverse populations 98%
- Extremely sparse models of linkage disequilibrium in ancestrally diverse association studies 98%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MUSSEL: Enhanced Bayesian Polygenic Risk Prediction Leveraging Information across Multiple Ancestry Groups 98%
- Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases 97%
- Incorporating family history of disease improves polygenic risk scores in diverse populations 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.