Back

JASMINE: A powerful representation learning method for enhanced analysis of incomplete multi-omics data

Ballard, J. L.; Dai, Z.; Shen, L.; Long, Q.

2025-06-22 bioinformatics
10.1101/2025.06.16.659949 bioRxiv
Show abstract

Integrative analysis of multi-omics data provides a more comprehensive and nuanced view of a subjects biological state. However, high-dimensionality and ubiquitous modality missingness present significant analytical challenges. Existing methods for incomplete multi-omics data are scarce, do not fully leverage both modality-specific and shared information, and produce task-biased representations. We propose JASMINE, a self-supervised representation learning method for incomplete multi-omics data that preserves both modality-specific and joint information and enhances sample similarity structure. JASMINE produces embeddings that achieve superior performance across multiple tasks for two different incomplete multi-omics datasets while requiring only a single round of training per dataset.

Published in npj Systems Biology and Applications · training set

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.