Back

STUltra: scalable and accurate integration for subcellular-level spatial omics data

Zhang, S.; Luo, S.; Luo, Y.; Su, S.; Liu, L.; Li, W.; Li, J.

2025-12-17 bioinformatics
10.64898/2025.12.15.694341 bioRxiv
Show abstract

Subcellular-level spatial transcriptomics data contain unprecedented contexts to uncover finer cellular clusters and their interactions. However, integrative analysis at subcellular resolution meets many challenging questions due to its ultra-large volume, ultra-high sparsity, and severe susceptibility to technical conditions and batch effects. We introduce STUltra, a scalable and accurate framework for integrating subcellular-level spatial omics data across spatial, temporal, and biomedical dimensions. Built on contrastive learning, STUltra combines a robust graph autoencoder with an interval sampling step to enhance batch-effect correction and enable clear characterization of shared and condition-specific tissue structures. It also provides seamless extension to super-resolution platforms such as Visium HD, Xenium, and Stereo-seq. STUltra is thus capable of identifying finer-grained cluster dynamics with distinguishable profile features, offering insights beyond those of prior studies. For example, STUltra successfully delineates interspersed macrophages within colorectal tumors, aligns mouse brain hippocampus across three subcellular platforms, and maps long-term muscle continuum during mouse embryonic development. Furthermore, from a mouse model of Alzheimers disease, STUltra detects disease-related astrocyte substructures and disentangles the regulatory network. Importantly, STUltra is remarkably scalable to process these datasets containing over 1,000,000 cells, outperforming existing tools in both accuracy and efficiency.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.