Scalable nonparametric clustering with unified marker gene selection for single-cell RNA-seq data
Nwizu, C.; Hughes, M.; Ramseier, M. L.; Navia, A.; Shalek, A. K.; Fusi, N.; Raghavan, S.; Winter, P. S.; Amini, A. P.; Crawford, L.
Show abstract
Clustering is commonly used in single-cell RNA-sequencing (scRNA-seq) pipelines to characterize cellular heterogeneity. However, current methods face two main limitations. First, they require user-specified heuristics which add time and complexity to bioinformatic workflows; second, they rely on post-selective differential expression analyses to identify marker genes driving cluster differences, which has been shown to be subject to inflated false discovery rates. We address these challenges by introducing nonparametric clustering of single-cell populations (NCLUSION): an infinite mixture model that leverages Bayesian sparse priors to identify marker genes while simultaneously performing clustering on single-cell expression data. NCLUSION uses a scalable variational inference algorithm to perform these analyses on datasets with up to millions of cells. Through simulations and analyses of publicly available scRNA-seq studies, we demonstrate that NCLUSION (i) matches the performance of other state-of-the-art clustering techniques with significantly reduced runtime and (ii) provides statistically robust and biologically relevant transcriptomic signatures for each of the clusters it identifies. Overall, NCLUSION represents a reliable hypothesis-generating tool for understanding patterns of expression variation present in single-cell populations.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- scBFA: modeling detection patterns to mitigate technical noise in large-scale single cell genomics data 97%
- A Systematic Evaluation of Single-cell RNA-sequencing Imputation Methods 97%
- Knowledge-primed neural networks enable biologically interpretable deep learning on single-cell sequencing data 97%
Similar papers in this journal
- HyGAnno: Hybrid graph neural network-based cell type annotation for single-cell ATAC sequencing data 96%
- CosGeneGate Selects Multi-functional and Credible Biomarkers for Single-cell Analysis 96%
- SnapHiC-G: identifying long-range enhancer-promoter interactions from single-cell Hi-C data via a global background model 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.