Normalization and gene selection for single-cell RNA-seq UMI data using sampling-adjusted sums of squares of Pearson residuals with a Poisson model
Klebanoff, V. F.
Show abstract
SCTransform in Seurat and scanpy.experimental.pp.recipe pearson residuals (scanpy henceforth) normalize UMI counts as Pearson residuals of negative binomial models. Residual variance scores genes for downstream analysis. Although we observed that both methods usually assign the highest scores to the same genes, for many highly ranked genes (e.g. among the top 2,000) scores may be unstable - not robust to the selection of cells used to calculate residuals. As an alternative, we consider the Poisson model, for which a natural score is the mean sum of squares of Pearson residuals. We show that these scores can be unstable if a genes nonzero UMI counts are concentrated on a small number of cells. This explains the instability for scanpy because of its similarity to the Poisson model. We define a metric for genes instability and observe that for all three methods it is negatively correlated with the number of cells on which genes counts are nonzero. To reduce the instability of scores based on the Poisson model, we score each gene using multiple random samples of approximately half of the cells. The minimum of these values defines a "sampling-adjusted" score. For data that we analyzed, these are more stable than scores from SCTransform and scanpy while generally agreeing with them on the highest ranked genes. As a second criterion to compare our proposal with SCTransform, we use differential expression analysis. For genes with high scores, the residuals Kruskal-Wallis H-statistics are generally greater for our method than for SCTransform and are more highly correlated with our methods scores.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 97%
- Generating Correlated Data for Omics Simulation 96%
- Binomial models uncover biological variation during feature selection of droplet-based single-cell RNA sequencing 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Gene prioritization based on random walks with restarts and absorbing states, to define gene sets regulating drug pharmacodynamics from single-cell analyses 94%
- Learning epistatic gene interactions from perturbation screens 94%
- Theoretical properties of nearest-neighbor distance distributions and novel metrics for high dimensional bioinformatics data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.