Back

Calculating sample size for identifying cell subpopulation in single-cell RNA-seq experiments

Kim, K. I.; Youn, A.; Bolisetty, M.; Palucka, K.; George, J.

2019-07-18 bioinformatics
10.1101/706481 bioRxiv
Show abstract

SO_SCPLOWUMMARYC_SCPLOWSingle-cell RNA sequencing (scRNA-seq) is a rapidly developing technology for studying gene expression at the individual cell level and is often used to identify subpopulations of cells. Although the use of scRNA-seq is steadily increasing in basic and translational research, there is currently no statistical model for calculating the optimal number of cells for use in experiments that seek to identify cell subpopulations. Here, we have developed a statistical method ncells for calculating the number of cells required to detect a rare subpopulation in a homogeneous cell population for the given type I and II error. ncells defines power as the probability of separation of subpopulations which is calculated from three user-defined parameters: the proportion of rare subpopulation, proportion of up-regulated marker genes of the subpopulation, and levels of differential expression of the marker genes. We applied ncells to the scRNA-seq data on dendritic cells and monocytes isolated from healthy blood donor to show its efficacy in calculating the optimal number of cells in identifying a novel subpopulation.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.