Back

MIND: Multimodal Integration with Neighbourhood-aware Distributions

Xing, H.; Yau, C.

2025-09-18 bioinformatics
10.1101/2025.09.15.676314 bioRxiv
Show abstract

Multi-omics profiling has become a powerful tool for biomedical applications such as cancer patient stratification and clustering. However, the characterisation and integration of multi-omics data remain challenging because of missingness and inherent heterogeneity. Methods such as imputation and sample exclusion often rely on strong assumptions that could potentially lead to information loss or distortion. To address these limitations, we propose a multi-omics integration framework that learns patient-specific embeddings from incomplete multiomics data based on a multimodal Variational Autoencoder with a data-driven prior. Specifically, we inject neighbourhood structure of the observed dataset encoded as affinity matrices into the prior of embeddings through exponential tilting, and use this prior to penalise the configuration of the latent embeddings based on the discrepancy between the neighbourhood structures in the data spaces and in the latent space. Our proposed method handles high missing rate and unbalanced missingness pattern well, and is robust in the presence of data with a low signal-to-noise ratio. Compared with existing data integration methods, the proposed method achieves better performance on a range of supervised and unsupervised downstreaming tasks on both synthetic and real data.

Published in Nature Communications (predicted rank #3) · training set

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.