Back

Unsupervised Deep Disentangled Representation of Single-Cell Omics

Moinfar, A. A.; Theis, F. J.

2025-08-19 bioinformatics
10.1101/2024.11.06.622266 bioRxiv
Show abstract

Single-cell genomics allows for the unbiased exploration of cellular heterogeneity. While deep generative models have become a keyprocessing step for integration and dimensionality reduction of single cell data, they function as black boxes that entangle biological signals and gene programs, restricting their interpretability and utility for downstream analysis. We address this limitation with Disentangled Representation Variational Inference (DRVI), a deep generative framework designed to provide dimension-wise interpretable disentangled representations for single-cell omics. DRVI isolates distinct sources of variation -- such as cell-type identity, developmental stage, cytokine signaling, and technical noise -- into individual, interpretable, and parallel factors without relying on supervised priors or restrictive linearity assumptions. This is achieved by exploiting additive decoders with a novel nonlinear pooling approach. We validate DRVIs capabilities to isolate cell types, secondary biological processes, perturbation effects, and stages in continuous trajectories on seven diverse datasets. In the Human Lung Cell Atlas, DRVI enables isolation of biological processes beyond cell types, covering interferon signaling, angiogenesis, proliferation, and stress response to metal ions, and identifies rare cell types such as migratory dendritic cells. In other examples, we show how to use interpretable factors to regress out isolated technical stress signals caused by sequencing, and to refine annotations for dendritic cells in an immune dataset. DRVIs interpretability does not come at the cost of integration performance, interpretable factors are consistent and transferable to similar datasets, and it generalizes easily to other modalities, such as single-cell chromatin accessibility. Altogether, DRVI provides a versatile framework for resolving the underlying structure of single-cell omics data.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.