Bayesian inference of the gene expression states of single cells from scRNA-seq data
Breda, J.; Zavolan, M.; van Nimwegen, E. J.
Show abstract
In spite of a large investment in the development of methodologies for analysis of single-cell RNA-seq data, there is still little agreement on how to best normalize such data, i.e. how to quantify gene expression states of single cells from such data. Starting from a few basic requirements such as that inferred expression states should correct for both intrinsic biological fluctuations and measurement noise, and that changes in expression state should be measured in terms of fold-changes rather than changes in absolute levels, we here derive a unique Bayesian procedure for normalizing single-cell RNA-seq data from first principles. Our implementation of this normalization procedure, called Sanity (SAmpling Noise corrected Inference of Transcription activitY), estimates log expression values and associated errors bars directly from raw UMI counts without any tunable parameters. Comparison of Sanity with other recent normalization methods on a selection of scRNA-seq datasets shows that Sanity outperforms other methods on basic downstream processing tasks such as clustering cells into subtypes and identification of differentially expressed genes. More importantly, we show that all other normalization methods present severely distorted pictures of the data. By failing to account for biological and technical Poisson noise, many methods systematically predict the lowest expressed genes to be most variable in expression, whereas in reality these genes provide least evidence of true biological variability. In addition, by confounding noise removal with lower-dimensional representation of the data, many methods introduce strong spurious correlations of expression levels with the total UMI count of each cell as well as spurious co-expression of genes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GoM DE: interpreting structure in sequence count data with differential expression analysis allowing for grades of membership 97%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 97%
- Visualizing scRNA-Seq Data at Population Scale with GloScope 97%
Similar papers in this journal
- mcRigor: a statistical method to enhance the rigor of metacell partitioning in single-cell data analysis 97%
- Flexible Experimental Designs for Valid Single-cell RNA-sequencing Experiments Allowing Batch Effects Correction 97%
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 96%
Similar papers in this journal
- TedSim: temporal dynamics simulation of single cell RNA-sequencing data and cell division history 96%
- Variance-adjusted Mahalanobis (VAM): a fast and accurate method for cell-specific gene set scoring 96%
- Epitome: Predicting epigenetic events in novel cell types with multi-cell deep ensemble learning 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.