Detecting Zero-Inflated Genes in Single-Cell Transcriptomics Data
Clivio, O.; Lopez, R.; Regier, J.; Gayoso, A.; Jordan, M. I.; Yosef, N.
Show abstract
In single-cell RNA sequencing data, biological processes or technical factors may induce an overabundance of zero measurements. Existing probabilistic approaches to interpreting these data either model all genes as zero-inflated, or none. But the overabundance of zeros might be gene-specific. Hence, we propose the AutoZI model, which, for each gene, places a spike-and-slab prior on a mixture assignment between a negative binomial (NB) component and a zero-inflated negative binomial (ZINB) component. We approximate the posterior distribution under this model using variational inference, and employ Bayesian decision theory to decide whether each gene is zero-inflated. On simulated data, AutoZI outperforms the alternatives. On negative control data, AutoZI retrieves predictions consistent to a previous study on ERCC spike-ins and recovers similar results on control RNAs. Applied to several datasets and instances of the 10x Chromium protocol, AutoZI allows both biological and technical interpretations of zero-inflation. Finally, AutoZIs decisions on mouse embyronic stem-cells suggest that zero-inflation might be due to transcriptional bursting.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- One model fits all: combining inference and simulation of gene regulatory networks 96%
- SCRaPL: hierarchical Bayesian modelling of associations in single cell multi-omics data 96%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 96%
Similar papers in this journal
- Benchmarking imputation methods for network inference using a novel method of synthetic scRNA-seq data generation 97%
- baredSC: Bayesian Approach to Retrieve Expression Distribution of Single-Cell 97%
- eSVD-DE: Cohort-wide differential expression in single-cell RNA-seq data using exponential-family embeddings 95%
Similar papers in this journal
- Cytomulate: accurate and efficient simulation of CyTOF data 96%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 96%
- Heterogeneous pseudobulk simulation enables realistic benchmarking of cell-type deconvolution methods 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.