Isolating salient variations of interest in single-cell transcriptomic data with contrastiveVI
Weinberger, E.; Lin, C.; Lee, S.-I.
Show abstract
Single-cell datasets are routinely collected to investigate changes in cellular state between control cells and corresponding cells in a treatment condition, such as exposure to a drug or infection by a pathogen. To better understand heterogeneity in treatment response, it is desirable to disentangle latent structures and variations uniquely enriched in treated cells from those shared with controls. However, standard computational models of single-cell data are not designed to explicitly separate these variations. Here, we introduce Contrastive Variational Inference (contrastiveVI; https://github.com/suinleelab/contrastiveVI), a framework for analyzing treatment-control scRNA-seq datasets that explicitly disentangles the data into shared and treatment-specific latent variables. Using four treatment-control scRNA-seq dataset pairs, we apply contrastiveVI to perform a broad set of standard analysis tasks, including visualization, clustering, and differential expression testing. In each case, we find that our method consistently achieves results that agree with known biological ground truths, while previously proposed methods often fail to do so. We conclude by generalizing our framework to multimodal measurements and applying it to analyze a single-cell dataset with joint transcriptome and surface protein measurements.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- ARTEMIS integrates autoencoders and schrodinger bridges to predict continuous dynamics of gene expression, cell population and perturbation from time-series single-cell data 97%
- Bayesian non-parametric clustering of single-cell mutation profiles 96%
- Single-cell copy number calling and event history reconstruction 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.