STELLAR: A flexible ensemble learning framework integrating rare variants to enhance polygenic risk prediction
Chen, T.; Li, X.; Mazumder, R.; Zhang, H.; Lin, X.
Show abstract
Whole-exome and whole-genome sequencing technology has enabled the discovery of rare genetic variants associated with human health and diseases. However, existing statistical methods used for rare variant association testing are not well-suited for building genetic risk prediction models that jointly incorporate rare and common variants. We propose STELLAR, a flexible ensemble learning-based approach to compute rare variant polygenic risk scores (PRS) using association summary statistics to enhance conventional common variant PRS. Our method combines burden-based and penalty-based rare variant analysis and leverages functional annotation information to prioritize potentially causal variants within the prediction models. In simulation studies, PRS using STELLAR consistently showed the highest prediction accuracy compared to models using common variants alone or rare variant burdens. Applied to UK Biobank whole-exome sequencing data (n=310,831) across eight continuous and five binary traits, STELLAR significantly improved prediction accuracy, refined stratification of individuals at the highest genetic risk beyond common variants, and prioritized biologically relevant genes. STELLAR provides a scalable strategy to incorporate rare variants into PRS in addition to common variants, advancing precision risk prediction and enabling more comprehensive assessment of genetic contributions to complex diseases.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 98%
- Combining case-control status and family history of disease increases association power 98%
- Scalable generalized linear mixed model for region-based association tests in large biobanks and cohorts 97%
Similar papers in this journal
- Incorporating family history of disease improves polygenic risk scores in diverse populations 98%
- Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases 97%
- Polygenic scores capture genetic modification of the adiposity-cardiometabolic risk factor relationship 97%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.