A Random Matrix Approach to Single Cell RNA-seq Analysis
Leviyang, S.
Show abstract
Single cell RNA-seq (scRNAseq) workflows typically start with a raw expression matrix and end with the clustering of sampled cells. Viewed broadly, scRNAseq is a signal processing workflow that takes a transcriptional signal as input and outputs a cell clustering. Currently, we lack a quantitative framework through which to describe the input signal and assess the dependence of correct clustering on the signal properties. As a result, fundamental questions regarding the resolution of scRNAseq remain unanswered and experimentalists have little guidance in determining whether a hypothesized cell type will be clustered by a particular scRNAseq experiment. In this work, we define the notion of a transcriptional signal associated with a gene module, show that the tools of random matrix theory can be used to characterize the signal as it moves through a common (PCA-based) scRNAseq workflow, and develop estimates for cell clustering based on the signal properties and, in particular, the signal strength. We give a formula - that can be computed from expression data - for the signal strength, providing a framework through which scRNAseq resolution can be investigated.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking imputation methods for network inference using a novel method of synthetic scRNA-seq data generation 96%
- eSVD-DE: Cohort-wide differential expression in single-cell RNA-seq data using exponential-family embeddings 96%
- baredSC: Bayesian Approach to Retrieve Expression Distribution of Single-Cell 96%
Similar papers in this journal
Similar papers in this journal
- Gene prioritization based on random walks with restarts and absorbing states, to define gene sets regulating drug pharmacodynamics from single-cell analyses 95%
- Partitioning gene-based variance of complex traits by gene score regression 94%
- Learning epistatic gene interactions from perturbation screens 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.