Back

Unifying Multimodal Single-Cell Data Using a Mixture of Experts β-Variational Autoencoder-Based Framework

Ashford, A. J.; Enright, T.; Nikolova, O.; Demir, E.

2025-03-06 bioinformatics
10.1101/2025.02.28.640429 bioRxiv
Show abstract

Multimodal single-cell assays measure complementary layers of cell state, but integration is complicated by differences in modality, sparsity, and cohort coverage. We present UniVI (Unified Variational Inference), a scalable mixture-of-experts {beta}-variational autoencoder that learns a shared latent space while preserving modality-specific structure. UniVI uses modality-specific encoders/decoders with a shared latent prior and a symmetric cross-modal alignment objective, enabling consistent integration of paired measurements without curated feature-link graphs or pre-annotated reference atlases; optional supervised heads can be added when labels are available. Across paired RNA-protein and RNA-chromatin datasets, UniVI yields coherent embeddings, improves label transfer, and supports cross-modal reconstruction and denoising. Extending to tri-modal measurements, UniVI maintains robust three-way alignment among RNA, chromatin accessibility, and surface proteins. UniVI also degrades smoothly under severe cell-type imbalance and in the presence of modality-exclusive populations. Finally, in an acute myeloid leukemia mosaic design, a paired RNA-protein bridge co-organizes independent RNA-only and protein+genotype cohorts, revealing genotype-associated neighborhoods that strengthen with mutation-head fine-tuning. Together, UniVI provides a flexible, interpretable framework for multimodal integration across paired, tri-modal, and mosaic study designs and supports practical reference-to-query projection in partially observed studies. The full package is available online at github.com/Ashford-A/UniVI.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.