Back

Enhancing Pan-cancer Spatial Transcriptomics atSingle-cell Resolution with stPainter

Yang, Y.; Luo, Y.; Zhang, K.; Zhang, Z.; Peng, H.; Cao, C.; Liu, Q.; Ma, B.; Chen, Y.; Shen, L.; Chen, E.

2026-02-13 bioinformatics
10.64898/2026.02.11.704553 bioRxiv
Show abstract

Subcellular spatial transcriptomics technologies offer unprecedented views of tissue architecture but are fundamentally constrained by sparse gene panels and limited detection sensitivity. Current computational enhancement strategies typically rely on tissue-matched single-cell RNA sequencing (scRNA-seq) references and necessitate computationally intensive retraining for each dataset, impeding their scalability and clinical applicability. Here, we present STPAINTER, a conditional generative model that leverages a massive pretraining pan-cancer scRNA-seq atlas to universally enhance spatial transcriptomics data. Built upon a latent diffusion architecture with stochastic differential equation-guided generation, STPAINTER learns a universal manifold of cellular states to reconstruct genome-wide expression profiles from sparse spatial measurements. Uniquely, our pretraining paradigm enables zero-shot generalization and empowers downstream tasks by providing imputed transcriptomes and informative latent variables to enhance resolution at both the gene and cluster levels. Applied STPAINTER upon 6 spatial transcriptomics datasets of different cancer types, we demonstrate that our model empowers high-fidelity downstream analyses, including fine-grained subpopulation clustering and pathway enrichment. Furthermore, cross-validation with spatially resolved proteomics (CODEX) confirms the biological veracity of the imputed cellular landscapes. STPAINTER provides a robust, scalable framework for decoding complex tumor microenvironments without the need for auxiliary sequencing data.

Published in Nature Communications (predicted rank #2) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.