Unsupervised learning of multi-omics data enables disease risk prediction in the UK Biobank
Rohrer, C.; Graef, J. F.; Pielies Avelli, M.; Medina, R. H.; Webel, H.; Ravn, K.; Rasmussen, S.
Show abstract
The size and complexity of biomedical datasets continue to grow, driving the development of methods that reduce dimensionality while preserving biological signals. Yet, when deep learning is applied to such data, the impact of preprocessing choices and dataset properties on model behavior is often overlooked. Here, we applied our framework Multi-Omics Variational autoEncoder (MOVE) to multiomics data from 452,026 UK Biobank participants, aiming to both evaluate the power of the learned representations for disease risk prediction and critically analyze how non-biological factors, like dataset properties and preprocessing decisions, can shape and influence the results. We show that reducing the dimensionality of the data by a factor of 80 still yields comparable prediction performance across 15 different diseases. We further demonstrate how dataset properties and preprocessing choices impact the model performance, latent representation and downstream results, and our findings strongly underline the need for thorough analysis and understanding of a models behavior before drawing conclusions from its results.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A tissue-aware machine learning framework enhances the mechanistic understanding and genetic diagnosis of Mendelian and rare diseases 96%
- Detection of PatIent-Level distances from single cell genomics and pathomics data with Optimal Transport (PILOT) 95%
- Causal integration of multi-omics data with prior knowledge to generate mechanistic hypotheses 94%
Similar papers in this journal
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 94%
- Conserved epigenetic regulatory logic infers genes governing cell identity 94%
- Multiomics and digital monitoring during lifestyle changes reveal independent dimensions of human biology and health 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.