A machine learning model for disease risk prediction by integrating genetic and non-genetic factors
Xu, Y.; Wang, C.; Li, Z.; Cai, Y.; Young, O.; Lyu, A.; Zhang, L.
Show abstract
Polygenic risk score (PRS) has been widely used to identify the high-risk individuals from the general population, which would be helpful for disease prevention and early treatment. Many methods have been developed to calculate PRS by weighted aggregating the phenotype-associated risk alleles from genome-wide association studies. However, only considering genetic effects may not be sufficient for risk prediction because the disease risk is not only related to genetic factors but also non-genetic factors, e.g., diet, physical exercise et al. But it is still a challenge to integrate these genetic and non-genetic factors into a unified machine learning framework for disease risk prediction. In this paper, we proposed PRSIMD (PRS Integrating Multi-source Data), a machine learning model that applies posterior regularization to integrate genetic and non-genetic factors to improve disease risk prediction. Also, we applied Mendelian Randomization analysis to identify the causal non-genetic risk factors for the selected diseases. We applied PRSIMD to predict type 2 diabetes and coronary artery disease from UK Biobank and observed that PRSIMD was significantly better than the methods to calculate PRS including p-value threshold (P+T), PRSice2, SBLUP, DMSLMM, and LDpred2. In addition, we observed that PRSIMD achieved the better predictive power than the composite risk score.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Disease Network Delineates the Disease Progression Profile of Cardiovascular Diseases 94%
- A Utility-Based Machine Learning-Driven Personalized Lifestyle Recommendation for Cardiovascular Disease Prevention 93%
- Developing deep learning-based strategies to predict the risk of hepatocellular carcinoma among patients with nonalcoholic fatty liver disease from electronic health records 93%
Similar papers in this journal
- Explainable deep transfer learning model for disease risk prediction using high-dimensional genomic data 96%
- Inferring latent temporal progression and regulatory networks from cross-sectional transcriptomic data of cancer samples 94%
- Sparse Multitask group Lasso for Genome-Wide Association Studies 93%
Similar papers in this journal
- CoMM-S2: a collaborative mixed model using summary statistics in transcriptome-wide association studies 94%
- High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning Models 94%
- DeepPerVar: a multimodal deep learning framework for functional interpretation of genetic variants in personal genome 94%
Similar papers in this journal
- PheCode-guided multi-modal topic modeling of electronic health records improves disease incidence prediction and GWAS discovery from UK Biobank 94%
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 93%
- CoRegNet: Unraveling Gene Co-regulation Networks from Public RNA-Seq Repositories Using a Beta-Binomial Statistical Model 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.