New algorithms for unsupervised cell clusteringfrom scRNA-seq data
Robles, M.; Diaz-Riano, J. I.; Forigua, C.; Ojeda, S.; Guio, L.; Siaucho, P.; Guzman-Porras, J.; Garcia-Orjuela, D.; Naranjo, A.; Maradei, S.; Quiroz, A.; Duitama, J.
Show abstract
The identification of cell types is a basic step of the pipeline for Single-Cell RNA sequencing data analysis. However, unsupervised clustering of cells from scRNA-seq data has multiple challenges: the high dimensional nature of the data, the sparse nature of the gene expression matrix, and the presence of technical noise that can introduce false zero entries. In this study, we introduce new algorithms for clustering scRNA-seq data. The first algorithm builds a k-MST graph from distances obtained directly from the input data without dimensionality reduction. The computation follows an iterative procedure of k steps in which each step calculates and stores the edges of minimum spanning trees over different subgraphs obtained removing edges selected in previous iterations. The Louvain algorithm is executed on the k-MST graph for cell clustering. We also explored alternatives based on neural networks in which an autoencoder is used to learn the parameters of a Gaussian mixture model, aiming to improve the handling of clusters with different shapes and sizes. Benchmark experiments with simulated data and public datasets show that the algorithms proposed in this work have competitive accuracy, compared to previous solutions, but also that sequencing depth, number of cells and tissue types have important effects on the performance of the algorithms. Moreover, we performed further experiments with scRNA-data taken from a patient with refractory epilepsy. The AE-GMM model achieved the best accuracy for this dataset, and the k-MST ranked first among methods that do not require previous information on the expected number of clusters.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 96%
- Coffee: Consensus Single Cell-Type Specific Inference For Gene Regulatory Networks 96%
- scaLR: a low-resource deep neural network-based platform for single cell analysis and biomarker discovery 95%
Similar papers in this journal
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 97%
- HiCImpute: A Bayesian Hierarchical Model for Identifying Structural Zeros and Enhancing Single Cell Hi-C Data. 95%
- G2S3: a gene graph-based imputation method for single-cell RNA sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.