ClustSIGNAL identifies cell types and subtypes using an adaptive smoothing approach for scalable spatial clustering
Panwar, P.; Guo, B.; Zhou, H.; Hicks, S. C.; Ghazanfar, S.
Show abstract
The increased uptake of high-resolution spatially-resolved transcriptomics (SRT) technologies demands the development of unsupervised methods to extract cell types and their spatial distribution from biological tissues. However, unsupervised clustering is challenging due to the sparsity of the data and the differences in cell arrangement within tissues. Here, we introduce ClustSIGNAL, a spatial clustering method that adaptively uses neighbourhood information to overcome data sparsity and perform cell type clustering. ClustSIGNAL first defines initial clusters and sub-clusters of cells with similar gene expression patterns. For each cell, a fixed neighbourhood size is defined, and entropy is calculated based on the proportion of initial subclusters in the neighbourhood to capture its composition. Cell-specific weights, generated from entropy values, are used to embed spatial information into the gene expression through adaptive smoothing. The transformed gene expression is then used for clustering cell types. We compared our adaptive smoothing approach with other smoothing scenarios on four simulated datasets of varying spatial complexity. We also evaluated our clustering method on four publicly available high-resolution SRT datasets and compared its performance to that of three other spatial clustering methods. We showed that ClustSIGNAL performs multi-sample clustering with high accuracy and can identify subtle cell types and subtypes of biological relevance. It is also robust to changes in spatial structure of tissues, segmentation errors, and sparsity. Overall, ClustSIGNAL stabilises gene expression of cells in homogeneous neighbourhoods and preserves distinct gene expression of cells in heterogeneous regions, effectively balancing the use of neighbouring cells as prior knowledge for downstream analysis. The ClustSIGNAL R/Bioconductor package is available from bioconductor.org/packages/clustSIGNAL.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Neighborhood nonnegative matrix factorization identifies patterns and spatially-variable genes in large-scale spatial transcriptomics data 97%
- scINSIGHT for interpreting single-cell gene expression from biologically heterogeneous data 96%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 96%
Similar papers in this journal
- Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST 97%
- Atlas-scale single-cell multi-sample multi-condition data integration using scMerge2 97%
- On the discovery of population-specific state transitions from multi-sample multi-condition single-cell RNA sequencing data 96%
Similar papers in this journal
- Randomized Spatial PCA (RASP): a computationally efficient method for dimensionality reduction of high-resolution spatial transcriptomics data 98%
- Non-linear Archetypal Analysis of Single-cell RNA-seq Data by Deep Autoencoders 96%
- Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics data 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.