Benchmarking Cell-Type-Specific Spatially Variable Gene Detection Methods Using a Realistic and Decomposable Simulation Framework
Li, W.; Ge, X.; Jiang, Y.
Show abstract
Identifying spatially variable genes within individual cell types is essential for characterizing spatially organized cell states and microenvironments from spatial transcriptomics data. Several computational methods have been developed for identifying cell-type-specific spatially variable genes (ctSVGs), but their relative performance and practical utility under realistic biological complexity remain largely unknown. To address this gap, we present the first systematic bench-mark study of all five existing ctSVG detection methods--CELINA, STANCE, C-SIDE, CTSV and spVC--using an integrated evaluation framework that combines idealized simulations, Xenium-based realistic simulations, and a decomposition-based diagnostic analysis. We compared the methods in terms of detection accuracy, scalability, and usability. Across realistic datasets generated on various tissue types, all methods experienced sharp declines in detection accuracy and substantial inflation of false discoveries compared to idealized simulations. To explain this failure, we developed a new simulation framework that decomposes the "realness" of the realistic simulation into interpretable biological and technical components, enabling us to attribute method-specific performance losses to specific components, including realistic diversity of cell types, heterogeneous cell layouts, null gene distributions, capture efficiency and realistic intra-cell-type spatial patterns. Together, our results show that no single method dominates across detection accuracy, scalability and usability, and we further clarify why current ctSVG methods fall short in realistic settings. We summarize these tradeoffs into a practical user guide to support method selection and highlight key challenges in developing robust, scalable ctSVG detection tools for real spatial transcriptomics data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BARcode DEmixing through Non-negative Spatial Regression (BarDensr) 95%
- Randomized Spatial PCA (RASP): a computationally efficient method for dimensionality reduction of high-resolution spatial transcriptomics data 95%
- Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics data 94%
Similar papers in this journal
- A comprehensive comparison on cell type composition inference for spatial transcriptomics data 96%
- BayeSMART: Bayesian Clustering of Multi-sample Spatially Resolved Transcriptomics Data 96%
- Graph Contrastive Learning of Subcellular-resolution Spatial Transcriptomics Improves Cell Type Annotation and Reveals Critical Molecular Pathways 95%
Similar papers in this journal
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 96%
- Heterogeneous pseudobulk simulation enables realistic benchmarking of cell-type deconvolution methods 96%
- Benchmarking algorithms for joint integration of unpaired and paired single-cell RNA-seq and ATAC-seq data 96%
Similar papers in this journal
- Belayer: Modeling discrete and continuous spatial variation in gene expression from spatially resolved transcriptomics 95%
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 95%
- An efficient not-only-linear correlation coefficient based on machine learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.