Probabilistic clustering of cells using single-cell RNA-seq data
Saha, J.; Tanvir, R. H.; Samee, M. A. H.; Rahman, A.
Show abstract
Single-cell RNA sequencing is a modern technology for analyzing cellular heterogeneity. A key challenge is to cluster a heterogeneous sample of different cell types into multiple different homogeneous groups. Although there exist a number of clustering methods, they do not perform well consistently across various datasets. Moreover, most of them are not based on probabilistic approaches making it difficult to assess uncertainties in their results. Therefore, in spite of having large cell atlases, it is often quite difficult to map cells to types. In addition, many of the methods require prior knowledge such as marker gene information for each type. Also due to technological limitations, dropouts of gene expressions may occur in the data which is not taken into account in other methods. Here we present a probabilistic method named CellHorizon for clustering scRNA-seq data that is based on a generative model, handles dropouts and works without any prior marker gene information. Experiments reveal that our method outperforms current state-of-the-art methods overall on six gold standard datasets.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 96%
- DGCyTOF: deep learning with graphic cluster visualization to predict cell types of single cell mass cytometry data 96%
- Assessing the Performance of Methods for Cell Clustering from Single-cell DNA Sequencing Data 96%
Similar papers in this journal
- SillyPutty: Improved clustering by optimizing the silhouette width 96%
- Differentially Expressed Heterogeneous Overdispersion Genes Testing for Count Data 95%
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.