Back

Comprehensive benchmarking of batch integration methods for spatial transcriptomics using a large-scale cancer atlas

Ouardini, K.; Ludington, L.; Loeb, R.; Secheresse, X.; Pignet, A.; Cabeli, V.; Domingues, O. D.

2026-01-13 cancer biology
10.64898/2026.01.12.699017 bioRxiv
Show abstract

Spatial transcriptomics (ST) enables spatially-resolved gene expression measurement, providing insights into tissue architecture and disease biology. However, batch effects from sequencing protocols, sample processing, and other technical factors can confound biological signals. Although batch correction has been extensively studied in single-cell transcriptomics, spatial integration methods lack rigorous benchmarking on large real-world datasets. This study benchmarks 11 representation-learning methods across three categories--linear, graph-based and probabilistic methods using Owkins MOSAIC Window dataset, a large-scale spatial transcriptomics atlas of human cancers. Methods are evaluated across three criteria: batch correction, biological conservation, and spatial conservation. We also propose a new integration metric to assess robustness of representations to domain shifts and generalizability to unseen samples. Probabilistic methods (scVIVA, scVI) outperform linear and graph-based approaches in batch correction and biological conservation. On the other hand Graph-based methods excelled at spatial conservation but underperformed in batch integration. Out-of-distribution evaluation reveals that sophisticated methods show reduced peformance on unseen samples while linear methods maintain robust generalization, highlighting trade-offs between integration quality and generalizability that should guide method selection for real-world applications.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.