Unified imputation of missing data modalities and features in multi-omic data via shared representation learning
Nambiar, A.; Melendez, C.; Noble, W. S.
Show abstract
Multi-omic studies promise a more comprehensive view of biological systems by jointly measuring multiple molecular layers. In practice, however, such datasets are rarely complete: entire molecular modalities may be missing for many samples, and observed modalities often contain substantial feature-level missingness. Existing imputation approaches typically address only one of these two problems, relying either on feature-level imputation within a single modality or on pairwise translation models that cannot accommodate arbitrary combinations of missing modalities. As a result, there is, to our knowledge, no unified framework for reconstructing both missing data modalities and missing values within those modalities. We present MIMIR, a deep learning framework for unified multi-omic imputation that addresses both missing modalities and missing values through shared representation learning. MIMIR first learns modality-specific representations using masked autoencoders and then projects these representations into a common latent space, enabling reconstruction from any subset of observed modalities. Evaluated on pan-cancer multi-omic data from The Cancer Genome Atlas, MIMIR consistently outperforms baseline methods across a range of missing-modality and missing-value scenarios, including missing completely at random and missing not at random settings. Analysis of the learned shared space reveals structured cross-modal dependencies that explain modality-specific differences in imputation accuracy, with transcriptional and epigenetic modalities forming a strongly aligned core and copy number variation contributing more distinct signal. Together, these results demonstrate that shared representation learning provides an effective and flexible foundation for multi-omic imputation under heterogeneous missingness. Availability and Implementationhttps://github.com/Noble-Lab/MIMIR
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scConfluence : single-cell diagonal integration with regularized Inverse Optimal Transport on weakly connected features 96%
- OmicVerse: A single pipeline for exploring the entire transcriptome universe 96%
- INSTINCT: Multi-sample integration of spatial chromatin accessibility sequencing data via stochastic domain translation 96%
Similar papers in this journal
Similar papers in this journal
- Multi-omics integration and regulatory inference for unpaired single-cell data with a graph-linked unified embedding framework 96%
- Density-Preserving Data Visualization Unveils Dynamic Patterns of Single-Cell Transcriptomic Variability 95%
- Dictionary learning for integrative, multimodal, and scalable single-cell analysis 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.