MultiPopPred: A Trans-Ethnic Disease Risk Prediction Method, and its Application to the South Asian Population
Kamal, R.; Narayanan, M.
Show abstract
Genome-wide association studies (GWAS) have guided significant contributions towards identifying disease associated Single Nucleotide Polymorphisms (SNPs) in Caucasian populations, albeit with limited focus on other understudied low-resource non-Caucasian populations. There have been active efforts over the years to understand and exploit the population specific versus shared aspects of the genotype-phenotype relation across different populations or ethnicities to bridge this gap. However, the efficacy of transfer learning models that are simpler than existing approaches and utilize individual-level data remains an open question. We propose MultiPopPred, a novel and simple trans-ethnic polygenic risk score (PRS) estimation method that taps into the shared genetic risk across populations and transfers information learned from multiple well-studied auxiliary populations to a less-studied target population. The default version of MultiPopPred (MPP-PRS+) harnesses individual-level data using a specially designed Nesterov-smoothed penalized shrinkage model and an L-BFGS optimization routine. Extensive comparative analyses performed on simulated genotype-phenotype data, assuming an infinitesimal model, reveal that MPP-PRS+ improves PRS prediction in the South Asian population by 38% on average across all simulation settings when compared to state-of-the-art trans-ethnic PRS estimation methods. This improvement is enhanced in settings with low target sample sizes and in semi-simulated settings. Furthermore, MPP-PRS+ produces better or comparable PRS predictions than state-of-the-art methods across 12 out of 16 evaluated quantitative and binary traits in UK Biobank, with the exception being 4 lipid-related traits. This performance trend is promising and encourages application of MultiPopPred for reliable PRS estimation in low-resource populations with individual-level data for complex omnigenic traits.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Haplotype based testing for a better understanding of the selective architecture 95%
- Detecting Selection in Low-Coverage High-Throughput Sequencing Data using Principal Component Analysis 95%
- Estimating the effect size of a hidden causal factor between SNPs and a continuous trait: a mediation model approach 94%
Similar papers in this journal
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 96%
- A novel haplotype-based eQTL approach identifies genetic associations not detected through conventional SNP-based methods 95%
- Efficient test for deviation from Hardy Weinberg Equilibrium with known or ambiguous typing in highly polymorphic loci 94%
Similar papers in this journal
- Sparse Multitask group Lasso for Genome-Wide Association Studies 97%
- Biological networks and GWAS: comparing and combining network methods to understand the genetics of familial breast cancer susceptibility in the GENESIS study 94%
- Association Tests Using Copy Number Profile Curves (CONCUR) Enhances Power in Rare Copy Number Variant Analysis 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.