Integrative machine learning approaches for predicting disease risk using multi-omics data from the UK Biobank
Aguilar, O. T.; Chang, C.; Bismuth, E.; Rivas, M. A.
Show abstract
We train prediction and survival models using multi-omics data for disease risk identification and stratification. Existing work on disease prediction focuses on risk analysis using datasets of individual data types (metabolomic, genomics, demographic), while our study creates an integrated model for disease risk assessment. We compare machine learning models such as Lasso Regression, Multi-Layer Perceptron, XG Boost, and ADA Boost to analyze multi-omics data, incorporating ROC-AUC score comparisons for various diseases and feature combinations. Additionally, we train Cox proportional hazard models for each disease to perform survival analysis. Although the integration of multi-omics data significantly improves risk prediction for 8 diseases, we find that the contribution of metabolomic data is marginal when compared to standard demographic, genetic, and biomarker features. Nonetheless, we see that metabolomics is a useful replacement for the standard biomarker panel when it is not readily available.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PaIRKAT: A pathway integrated regression-based kernel association test with applications to metabolomics and COPD phenotypes 93%
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 93%
- Using random forests to uncover the predictive power of distance-varying cell interactions in tumor microenvironments 93%
Similar papers in this journal
- High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning Models 94%
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 94%
- Joint Modeling of Longitudinal Biomarker and Survival Outcomes with the Presence of Competing Risk in Nested Case-Control Studies with Application to the TEDDY Microbiome Dataset 94%
Similar papers in this journal
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 94%
- Biobank-scale methods and projections for sparse polygenic prediction from machine learning 94%
- Polygenic Health Index, General Health, Pleiotropy, Embryo Selection and Disease Risk 94%
Similar papers in this journal
Similar papers in this journal
- A Regularized Cox Hierarchical Model for Incorporating Annotation Information in Predictive Omic Studies 94%
- A compact encoding of the genome suitable for machine learning prediction of traits and genetic risk scores. 93%
- Expanding a Database-derived Biomedical Knowledge Graph via Multi-relation Extraction from Biomedical Abstracts 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.