Back

Scalable integration and prediction of unpaired single-cell and spatial multi-omics via regularized disentanglement

Sun, J.; Liang, C.; Wei, R.; Zheng, P.; Yan, H.; Bai, L.; Zhang, K.; Ouyang, W.; Ye, P.

2025-11-29 genomics
10.1101/2025.11.25.689803 bioRxiv
Show abstract

Understanding cellular states urgently requires methods capable of integrating large-scale, heterogeneous single-cell and spatial omics data. However, these data are often completely unpaired due to destructive assays and suffer from technical noise, variable feature coverage, and immense scale. We present scMRDR, a scalable computational framework leveraging regularized disentangled representation learning to integrate multiple, completely unpaired single-cell omics datasets with heterogeneous resolutions and coverages. scMRDR overcomes common data-pairing requirements and computational bottlenecks by learning a unified, structure-preserving latent embedding, efficiently scalable to large-scale multi-omics data. This integrated representation further enables robust cross-modal translation like predicting chromatin accessibility from gene expression and, critically, allows for the imputation of spatial coordinates onto non-spatial single-cell modalities using a reference atlas. This spatial mapping capability provides the necessary input for sophisticated, spatially-aware statistical models, enabling the identification of novel spatially variable genes and the dissection of epigenetic regulatory programs within their native tissue context.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.