Back

STCC: consensus clustering enhances spatial domain detection for spatial transcriptomics data

Hu, C.; Wei, N.; Yang, J.; Wu, H.-J.; Zheng, X.

2024-02-28 bioinformatics
10.1101/2024.02.25.581996 bioRxiv
Show abstract

The rapid advance of spatially resolved transcriptomics technologies has yielded substantial spatial transcriptomics data. Deriving biological insights from these data poses non-trivial computational and analysis challenges, of which the most fundamental step is spatial domain detection (or spatial clustering). Although a number of tools for spatial domain detection have been proposed in recent years, their performance varies across datasets and experimental platforms. It is thus an important task to take full advantage of different tools to get a more accurate and stable result through consensus strategy. In this work, we developed STCC, a novel consensus clustering framework for spatial transcriptomics data that aggregates outcomes from state-of-the-art tools using a variety of consensus strategies, including Onehot-based, Average-based, Hypergraph-based and wNMF-based methods. Comprehensive assessments on simulated and real data from distinct experimental platforms show that consensus clustering significantly improves clustering accuracy over individual methods under varied input parameters. For normal tissue samples exhibiting clear layered structure, consensus clustering by integrating multiple baseline methods leads to improved results. Conversely, when analyzing tumor samples that display scattered cell type distribution patterns, integration of a single baseline method yields satisfactory performance. For consensus strategies, Average-based and Hypergraph-based approaches demonstrated optimal precision and stability. Overall, STCC provides a scalable and practical solution for spatial domain detection in spatial transcriptomic data, laying a solid foundation for future research and applications in spatial transcriptomics.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.