SPLENDID incorporates continuous genetic ancestry in biobank-scale data to improve polygenic risk prediction across diverse populations
Chen, T.; Zhang, H.; Mazumder, R.; Lin, X.
Show abstract
Polygenic risk scores are widely used in disease risk stratification, but their accuracy varies across diverse populations. Recent methods large-scale leverage multi-ancestry data to improve accuracy in under-represented populations but require labelling individuals by ancestry for prediction. This poses challenges for practical use, as clinical practices are typically not based on ancestry. We propose SPLENDID, a novel penalized regression framework for diverse biobank-scale data. Our method utilizes ancestry principal component interactions to model genetic ancestry as a continuum within a single prediction model for all ancestries, eliminating the need for discrete labels. In extensive simulations and analyses of 9 traits from the All of Us Research Program (N=224,364) and UK Biobank (N=340,140), SPLENDID significantly outperformed existing methods in prediction accuracy and model sparsity. By directly incorporating continuous genetic ancestry in model training, SPLENDID stands as a valuable tool for robust risk prediction across diverse populations and fairer clinical implementation.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 98%
- Combining case-control status and family history of disease increases association power 98%
- Leveraging fine-mapping and non-European training data to improve trans-ethnic polygenic risk scores 98%
Similar papers in this journal
- CADET: Enhanced transcriptome-wide association analyses in admixed samples using eQTL summary data 98%
- Enrichment analyses identify shared associations for 25 quantitative traits in over 600,000 individuals from seven diverse ancestries 97%
- Characterizing substructure via mixture modeling in large-scale genetic summary statistics 97%
Similar papers in this journal
- Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases 98%
- Incorporating family history of disease improves polygenic risk scores in diverse populations 98%
- Polygenic scores capture genetic modification of the adiposity-cardiometabolic risk factor relationship 97%
Similar papers in this journal
- Quantifying portable genetic effects and improving cross-ancestry genetic prediction with GWAS summary statistics 98%
- Leveraging information between multiple population groups and traits improves fine-mapping resolution 97%
- Population-specific causal disease effect sizes in functionally important regions impacted by selection 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.