PCQC: Selecting optimal principal components for identifying clusters with highly imbalanced class sizes in single-cell RNA-seq data
Burstein, D.; Fullard, J.; Roussos, P.
Show abstract
SummaryPrior to identifying clusters in single cell gene expression experiments, selecting the top principal components is a critical step for filtering out noise in the data set. Identifying these top principal components typically focuses on the total variance explained, and principal components that explain small clusters from rare populations will not necessarily capture a large percentage of variance in the data. We present a computationally efficient alternative for identifying the optimal principal components based on the tails of the distribution of variance explained for each observation. We then evaluate the efficacy of our approach in three different single cell RNA-sequencing data sets and find that our method matches, or outperforms, other selection criteria that are typically employed in the literature. Availability and implementationpcqc is written in Python and available at github.com/RoussosLab/pcqc
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scConsensus: combining supervised and unsupervised clustering for cell type identification in single-cell RNA sequencing data 95%
- CDSeqR: fast complete deconvolution for gene expression data from bulk tissues 95%
- GEOlimma: Differential Expression Analysis and Feature Selection Using Pre-Existing Microarray Data 94%
Similar papers in this journal
Similar papers in this journal
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 96%
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 96%
- STREAK: A Supervised Cell Surface Receptor Abundance Estimation Strategy for Single Cell RNA-Sequencing Data using Feature Selection and Thresholded Gene Set Scoring 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.