A unified model for interpretable latent embedding of multi-sample, multi-condition single-cell data
Madrigal, A.; Lu, T.; Soto, L. M.; Najafabadi, H. S.
Show abstract
Analysis of single cells across multiple samples and/or conditions encompasses a series of interrelated tasks, which range from normalization and inter-sample harmonization to identification of cell state shifts associated with experimental conditions. Other downstream analyses are further needed to annotate cell states, extract pathway-level activity metrics, and/or nominate gene regulatory drivers of cell-to-cell variability or cell state shifts. Existing methods address these analytical requirements sequentially, lacking a cohesive framework to unify them. Moreover, these analyses are currently confined to specific modalities where the biological quantity of interest gives rise to a singular measurement. However, other modalities require joint consideration of dual measurements; for example, modeling the latent space of alternative splicing involves joint analysis of exon inclusion and exclusion reads. Here, we introduce a generative model, called GEDI, to identify latent space variations in multi-sample, multi-condition single cell datasets and attribute them to sample-level covariates. GEDI enables cross-sample cell state mapping on par with the state-of-the-art integration methods, cluster-free differential gene expression analysis along the continuum of cell states in the form of transcriptomic vector fields, and machine learning-based prediction of sample characteristics from single-cell data. By incorporating gene-level prior knowledge, it can further project pathway and regulatory network activities onto the cellular state space, enabling the computation of the gradient fields of transcription factor activities and their association with the transcriptomic vector fields of sample covariates. Finally, we demonstrate that GEDI surpasses the gene-centric approach by extending all these concepts to the study of alternative cassette exon splicing and mRNA stability landscapes in single cells.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- OmicVerse: A single pipeline for exploring the entire transcriptome universe 98%
- Pathway Centric Analysis for single-cell RNA-seq and Spatial Transcriptomics Data with GSDensity 98%
- scConfluence : single-cell diagonal integration with regularized Inverse Optimal Transport on weakly connected features 97%
Similar papers in this journal
- scAlign: a tool for alignment, integration and rare cell identification from scRNA-seq data 98%
- Smoother: A Unified and Modular Framework for Incorporating Structural Dependency in Spatial Omics Data 97%
- geneBasis: an iterative approach for unsupervised selection of targeted gene panels from scRNA-seq. 97%
Similar papers in this journal
- scTrace+: enhance the cell fate inference by integrating the lineage-tracing and multi-faceted transcriptomic similarity information 96%
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 96%
- scCausalVI disentangles single-cell perturbation responses with causality-aware generative model 96%
Similar papers in this journal
- Probabilistic tensor decomposition extracts better latent embeddings from single-cell multiomic data 97%
- Interpretable trajectory inference with single-cell Linear Adaptive Negative-binomial Expression (scLANE) testing 97%
- Coralysis enables sensitive identification of imbalanced cell types and states in single-cell data via multi-level integration 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.