Applying Machine Learning on UK Biobank biomarker data empowers case-control discovery yield
Garg, M.; Karpinski, M.; Matelska, D.; Middleton, L.; Mitchell, J.; O'Neill, A.; Wang, Q.; Harper, A. R.; Dhindsa, R. S.; Petrovski, S.; Vitsios, D.
Show abstract
Missing or inaccurate diagnoses in biobank datasets can reduce the power of human genetic association studies. We present a machine-learning framework (MILTON) that utilizes the wealth of phenotypic information available in a biobank dataset to identify undiagnosed individuals within the cohort who have biomarker profiles similar to those of positively diagnosed cases. We applied MILTON to perform an augmented phenome-wide association study (PheWAS) based on 405,703 whole exome sequencing samples from UK Biobank, resulting in improved signals for known (p<1x10-8) gene-disease relationships alongside 206 novel gene-disease relationships that only achieved genome-wide significance upon using MILTON. To further validate these putatively novel discoveries, we adopt two orthogonal machine learning methods that prioritise gene-disease relationships using comprehensive publicly available datasets alongside a biological insights knowledge graph. For additional clinical translation utility, MILTON outputs a disease-specific biomarker set per disease as well as comorbidity clusters across ICD10 disease codes based on shared biomarker profiles of positively labelled cases. All the extracted associations and biomarker importance results for the 3,308 studied binary traits will be made available via an interactive web-portal.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Phenome-wide Mendelian randomization mapping the influence of the plasma proteome on complex diseases 96%
- Systematic assessment of regulatory effects of human disease variants in pluripotent cells 96%
- Central role of glycosylation processes in human genetic susceptibility to SARS-CoV-2 infections with Omicron variants 96%
Similar papers in this journal
- Proteome-wide Mendelian randomization in global biobank meta-analysis reveals multi-ancestry drug targets for common diseases 96%
- Incorporating family history of disease improves polygenic risk scores in diverse populations 96%
- The functional impact of rare variation across the regulatory cascade 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.