Improved prediction of blood biomarkers using deep learning
Sigurdsson, A. I.; Ravn, K.; Winther, O.; Lund, O.; Brunak, S.; Vilhjalmsson, B. J.; Rasmussen, S.
Show abstract
Blood and urine biomarkers are an essential part of modern medicine, not only for diagnosis, but also for their direct influence on disease. Many biomarkers have a genetic component, and they have been studied extensively with genome-wide association studies (GWAS) and methods that compute polygenic scores (PGSs). However, these methods generally assume both an additive allelic model and an additive genetic architecture for the target outcome, and thereby risk not capturing non-linear allelic effects nor epistatic interactions. Here, we trained and evaluated deep-learning (DL) models for PGS prediction of 34 blood and urine biomarkers in the UK Biobank cohort, and compared them to linear methods. For lipid traits, the DL models greatly outperformed the linear methods, which we found to be consistent across diverse populations. Furthermore, the DL models captured non-linear effects in covariates, non-additive genotype (allelic) effects, and epistatic interactions between SNPs. Finally, when using only genome-wide significant SNPs from GWAS, the DL models performed equally well or better for all 34 traits tested. Our findings suggest that DL can serve as a valuable addition to existing methods for genotype-phenotype modelling in the era of increasing data availability.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genetic analysis of blood molecular phenotypes reveals regulatory networks affecting complex traits: a DIRECT study 96%
- Testing and controlling for horizontal pleiotropy with the probabilistic Mendelian randomization in transcriptome-wide association studies 96%
- Uncovering causal gene-tissue pairs and variants: A multivariable TWAS method controlling for infinitesimal effects 96%
Similar papers in this journal
- Polygenic scores capture genetic modification of the adiposity-cardiometabolic risk factor relationship 97%
- Characterizing the genetic architecture of drug response using gene-context interaction methods 96%
- Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases 96%
Similar papers in this journal
- Calibrated prediction intervals for polygenic scores across diverse contexts 97%
- Leveraging fine-mapping and non-European training data to improve trans-ethnic polygenic risk scores 96%
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.