Designing Single-Cell RNA-Sequencing Experiments for Learning Latent Representations
Treppner, M.; Haug, S.; Köttgen, A.; Binder, H.
Show abstract
To investigate the complexity arising from single-cell RNA-sequencing (scRNA-seq) data, researchers increasingly resort to deep generative models, specifically variational autoencoders (VAEs), which are trained by variational inference techniques. Similar to other dimension reduction approaches, this allows encoding the inherent biological signals of gene expression data, such as pathways or gene programs, into lower-dimensional latent representations. However, the number of cells necessary to adequately uncover such latent representations is often unknown. Therefore, we propose a single-cell variational inference approach for designing experiments (scVIDE) to determine statistical power for detecting cell group structure in a lower-dimensional representation. The approach is based on a test statistic that quantifies the contribution of every single cell to the latent representation. Using a smaller scRNA-seq data set as a starting point, we generate synthetic data sets of various sizes from a fitted VAE. Employing a permutation technique for obtaining a null distribution of the test statistic, we subsequently determine the statistical power for various numbers of cells, thus guiding experimental design. We illustrate with several data sets from various sequencing protocols how researchers can use scVIDE to determine the statistical power for cell group detection within their own scRNA-seq studies. We also consider the setting of transcriptomics studies with large numbers of cells, where scVIDE can be used to determine the statistical power for sub-clustering. For this purpose, we use data from the human KPMP Kidney Cell Atlas and evaluate the power for sub-clustering of the epithelial cells contained therein. To make our approach readily accessible, we provide a comprehensive Jupyter notebook at https://github.com/MTreppner/scVIDE.jl that researchers can use to design their own experiments based on scVIDE.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- ACTIVA: realistic single-cell RNA-seq generation with automatic cell-type identification using introspective variational autoencoders 97%
- Identifying cancer pathway dysregulations using differential causal effects 97%
- BEATRICE: Bayesian Fine-mapping from Summary Datausing Deep Variational Inference 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Synthetic observations from deep generative models and binary omics data with limited sample size 97%
- Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing data 97%
- Graph Contrastive Learning as a Versatile Foundation for Advanced scRNA-seq Data Analysis 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.