Back

Towards a Better Understanding of Batch Effects in Spatial Transcriptomics: Definition and Method Evaluation

Zhang, Y.; Hou, Q.

2025-03-13 bioinformatics
10.1101/2025.03.12.642755 bioRxiv
Show abstract

BackgroundSpatial transcriptomics (ST) enables high-resolution mapping of gene expression within tissue slices, providing detailed insights into tissue architecture and cellular interactions. However, batch effects, arising from non-biological variations in sample collection, processing, sequencing platforms, or experimental protocols, can obscure biological signals, hinder data integration, and impact downstream analyses. Despite their critical impact, batch effects in ST datasets remain poorly defined and insufficiently explored. To address this gap, we propose a framework to categorize and define batch effects in ST and systematically evaluate the performance of ST methods with batch effect correction capabilities. ResultsWe categorized batch effects in ST into four types based on their sources: (1) Inter-slice, (2) Inter-sample, (3) Cross-protocol/platform, and (4) Intra-slice. Seven ST integration methods--DeepST, STAligner, GraphST, STitch3D, PRECAST, spatiAlign, and SPIRAL--were evaluated on benchmark datasets from human and mouse tissues. Using metrics such as graph connectivity, kBET, ASW, and iLISI, we assessed both the preservation of biological neighborhoods and the effectiveness of these methods in batch correction. Additionally, we applied STAligner for downstream analysis to compare results before and after batch correction, further highlighting the importance of batch effect correction in ST analysis. ConclusionNo single method is universally optimal. GraphST, PRECAST, SPIRAL, and STAligner performed well for same-platform integration, whereas SPIRAL and STAligner excelled in cross-platform settings. These findings highlight the need for robust and generalizable ST approaches with effective batch correction capabilities to facilitate the integration of multi-platform ST datasets in future research.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.