Back

Tissue tearing degrades optimal-transport and diffeomorphic registration of spatial transcriptomics beyond displacement magnitude: a multi-seed deformation benchmark and a supervised graph cross-attention proof-of-concept.

Maniar, R. K.; Lee, S. G.; Lee, S. S.

2026-07-05 bioinformatics
10.64898/2026.06.30.735390 bioRxiv
Show abstract

Background. Three-dimensional reconstruction from serial spatial-transcriptomics (ST) sections requires registering adjacent slices, but physical sectioning introduces tears -- discontinuous, non-isometric deformations. Leading methods rely on priors that tears strain: PASTE/PASTE2 use Fused Gromov-Wasserstein optimal transport (OT), which assumes near-isometric preservation of within-slice distances, while STalign and CODA use diffeomorphic (LDDMM) mapping, which cannot change tissue topology. Learned-deformation ST methods are emerging (STaCker, INST-Align), but OT/diffeomorphic behaviour under tearing has not been systematically characterised. Methods. On the spatialLIBD human DLPFC Visium dataset (Maynard et al., 2021; 3 donors), we build a controlled benchmark -- known smooth warps, single-block rigid tears (expression unchanged), and an identity self-control -- at severities of 0-8 spot pitches, scored against an approximate array-position ground truth (~8 px residual). We evaluate three unsupervised incumbents -- PASTE2 (OT, over five warp seeds), STalign (diffeomorphic LDDMM), and GPSA (Gaussian-process warp) -- add a magnitude-matched smooth control, and test a minimal graph model, Sutura (per-slice graph encoder -> cross-attention correspondence -> per-spot displacement; spatial coupling is local kNN message passing only, no explicit smoothness penalty). Sutura is trained supervised on each tissue's ground truth; all baselines are unsupervised. Generalisation is assessed by leave-one-donor-out across all three donors. Results. OT registration is robust to smooth warps but degrades reproducibly under tearing: nearest-correspondence (argmax) error 722 +/- 5 -> 855 +/- 27 px and layer accuracy 64.9% -> 60.5% (mean +/- 95% CI, 5 seeds). The effect is not merely displacement magnitude: at a matched mean displacement (~2000 px), a smooth warp costs 769 px / 60.2% accuracy whereas a tear costs 863 px / 57.5% -- an extra ~100 px and ~3 points attributable to the discontinuity. STalign (LDDMM) and GPSA (GP warp) both collapse at severe tears (866 px and 931 px respectively), confirming tear-collapse is field-wide across three independent method families. Trained and evaluated on the same donor, Sutura fits torn-tissue correspondence to a median 99 -> 106 px (5-seed), but under leave-one-donor-out is 1236 +/- 2 -> 1584 +/- 52 px -- approximately 1.8-3.6x worse than PASTE2 on every unseen donor. A contrastive correspondence loss halves the gap on two of three donors (to 816 -> 949 and 749 -> 826 px, approximately 1.1-1.2x PASTE2 at worst-case tear) but is modest on the third and never surpasses PASTE2. Conclusion. Tearing is a real, magnitude-controlled failure mode of all three incumbent method classes. A learned model fits it in-sample but donor-invariant generalisation remains open. The contrastive fix roughly halves the held-out gap on two of three donors and nears PASTE2 at worst-case tear, but does not surpass it: donor-invariance is improved, not solved. The durable contribution is the benchmark, the characterisation across three method families, and an honest negative with a diagnosed mechanism.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Nature Methods
385 papers in training set
Top 0.6%
12.9%
2
Briefings in Bioinformatics
354 papers in training set
Top 0.4%
12.1%
3
Bioinformatics
1204 papers in training set
Top 3%
8.0%
4
Nature Communications
5641 papers in training set
Top 20%
8.0%
5
Scientific Reports
3612 papers in training set
Top 26%
4.1%
6
Medical Image Analysis
35 papers in training set
Top 0.2%
3.3%
7
PLOS ONE
5266 papers in training set
Top 37%
3.3%
50% of probability mass above
8
Nature Machine Intelligence
70 papers in training set
Top 0.9%
3.2%
9
Genome Biology
637 papers in training set
Top 4%
2.8%
10
PLOS Computational Biology
1863 papers in training set
Top 11%
2.7%
11
GigaScience
212 papers in training set
Top 1%
2.7%
12
eLife
5828 papers in training set
Top 40%
2.5%
13
Nature
645 papers in training set
Top 5%
2.5%
14
Human Brain Mapping
329 papers in training set
Top 2%
2.4%
15
npj Digital Medicine
118 papers in training set
Top 2%
1.9%
16
Communications Biology
993 papers in training set
Top 12%
1.9%
17
Cell Systems
201 papers in training set
Top 3%
1.1%
18
Nature Biomedical Engineering
47 papers in training set
Top 1.0%
1.1%
19
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 37%
1.1%
20
Nucleic Acids Research
1281 papers in training set
Top 12%
1.1%
21
Nature Biotechnology
172 papers in training set
Top 4%
0.9%
22
NeuroImage
903 papers in training set
Top 6%
0.9%
23
Bioinformatics Advances
203 papers in training set
Top 4%
0.9%
24
iScience
1154 papers in training set
Top 33%
0.9%
25
Patterns
78 papers in training set
Top 3%
0.9%
26
Scientific Data
209 papers in training set
Top 3%
0.6%
27
Cell Reports Methods
165 papers in training set
Top 4%
0.6%
28
New Phytologist
346 papers in training set
Top 5%
0.6%