A Bayesian nonparametric semi-supervised model for integration of multiple single-cell experiments
Verma, A.; Engelhardt, B. E.
Show abstract
Joint analysis of multiple single cell RNA-sequencing (scRNA-seq) data is confounded by technical batch effects across experiments, biological or environmental variability across cells, and different capture processes across sequencing platforms. Manifold alignment is a principled, effective tool for integrating multiple data sets and controlling for confounding factors. We demonstrate that the semi-supervised t-distributed Gaussian process latent variable model (sstGPLVM), which projects the data onto a mixture of fixed and latent dimensions, can learn a unified low-dimensional embedding for multiple single cell experiments with minimal assumptions. We show the efficacy of the model as compared with state-of-the-art methods for single cell data integration on simulated data, pancreas cells from four sequencing technologies, induced pluripotent stem cells from male and female donors, and mouse brain cells from both spatial seqFISH+ and traditional scRNA-seq. Code and data is available at https://github.com/architverma1/sc-manifold-alignment
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Scalable Integration of Multiomic Single Cell Data Using Generative Adversarial Networks 97%
- SpatialRNA: a Python package for easy application of Graph Neural Network models on single-molecule spatial transcriptomicsdataset 95%
- FateNet: an integration of dynamical systems and deep learning for cell fate prediction 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- scDREAMER: atlas-level integration of single-cell datasets using deep generative model paired with adversarial classifier 96%
- Deep generative model embedding of single-cell RNA-Seq profiles on hyperspheres and hyperbolic spaces 95%
- mcRigor: a statistical method to enhance the rigor of metacell partitioning in single-cell data analysis 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.