Comparison of Dimensionality Reduction and Clustering Methods for Single-Cell Transcriptomics Data
Pandit, V.; Swain, A. K.; Yadav, P.
Show abstract
Dimensionality reduction (DR) methods are applied to extract relevant features from inherently high dimensional and noisy single-cell RNA sequencing (scRNA-seq) data. Choice of DR method could influence the performance of clustering algorithm and subsequent analysis outcomes. We performed a benchmarking study of seven popular DR methods and four clustering algorithms widely used for scRNA-seq datasets. For this purpose, we used three publicly available real scRNA-seq datasets. The performance was evaluated using two clustering metrics viz. adjusted random index (ARI) and normalized mutual index (NMI). We also compared our results with a similar study published by Xiang and colleagues. Overall, we observed higher ARI and NMI scores for DR methods when compared with Xiangs study. We also noticed several differences between our and Xiangs study. Noteworthy, three methods, namely, Independent Component Analysis (ICA), t-Distributed Stochastic Neighbor Embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP) performed consistently well across three datasets. Linear method ICA was best performer on Segerstolpe dataset, while nonlinear methods UMAP and t-SNE best performed on Deng and Chu datasets, respectively. Neural network-based methods Variational Autoencoder (VAE) and Deep Count Autoencoder (DCA) could not perform well probably due to their sensitivity to hyperparameters and overfitting. Among clustering methods, Gaussian Mixture Models (GMMs) performed consistently well across datasets. This might be because GMMs are the universal approximators of posterior probability densities. We conclude that performance of different DR methods is more dataset dependent and for various scRNA-seq datasets different algorithms are more suited and there is no one-fit-all method.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 96%
- SillyPutty: Improved clustering by optimizing the silhouette width 95%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 95%
Similar papers in this journal
- Normalization of RNA-Seq Data using Adaptive Trimmed Mean with Multi-reference 96%
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 96%
- Feature Extraction Approaches for Biological Sequences: A Comparative Study of Mathematical Models 96%
Similar papers in this journal
- A Robust Spike Sorting Method based on the Joint Optimization of Linear Discrimination Analysis and Density Peaks 96%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 95%
- Tensor decomposition- and principal component analysis-based unsupervised feature extraction to select more reasonable differentially expressed genes: Optimization of standard deviation versus state-of-art methods 95%
Similar papers in this journal
- VICTOR: A visual analytics web application for comparing cluster sets 96%
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 96%
- Enrichment analysis on regulatory subspaces: a novel direction for the superior description of cellular responses to SARS-CoV-2 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.