OT-knn: a neighborhood-aware optimal transport framework for aligning spatial transcriptomics data
Song, J.; Li, Q.
Show abstract
Spatial transcriptomics (ST) measures gene expression while preserving spatial context within tissues, enabling detailed characterization of tissue organization. As ST technologies advance, aligning datasets across tissue sections, individuals, platforms, and developmental stages has become increasingly important but remains challenging due to sparse expression, biological heterogeneity, and geometric distortions between slices. We introduce OT-knn, a method for ST alignment that integrates local neighborhood information within an optimal transport framework. Rather than relying solely on single-spot expression, OT-knn reconstructs each spot using its spatial k-nearest neighbors, capturing microenvironment context that is more robust to noise and variability. These representations are then used to derive probabilistic correspondences between slices. We evaluate OT-knn using simulated data with known ground-truth alignment and real datasets from multiple ST platforms, including human dorsolateral prefrontal cortex data (10x Genomics Visium), mouse brain aging data with both within-donor and cross-donor comparisons (MERFISH), and a multi-stage axolotl brain dataset (Stereo-seq). Across these settings, OT-knn achieves accurate and robust alignment, particularly in the presence of spatial deformation, donor heterogeneity, and developmental variation. Author summaryUnderstanding how cells are arranged within tissues is important for studying how organs develop, age, and respond to disease. Spatial transcriptomics is a new technology that measures gene activity while keeping track of where each measurement comes from in the tissue. However, comparing data from different tissue slices, different individuals, or different time points is difficult. Tissues can change shape during preparation, gene measurements can be noisy, and natural biological differences exist across samples. In this work, we develop a method to better match corresponding regions across tissue slices. Instead of looking at each location by itself, we also consider information from its surrounding neighborhood. This provides a more stable and informative view of the tissue and makes it easier to match similar regions, even when the data are imperfect or the tissue shapes differ. We test our method on both simulated and real datasets from human, mouse, and axolotl brain tissues. Our approach enables more reliable comparisons of spatial gene activity across samples, which can help researchers study development, aging, and disease.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- stDyer enables spatial domain clustering with dynamic graph embedding 97%
- Cross-species imputation and comparison of single-cell transcriptomic profiles 97%
- scCross: A Deep Generative Model for Unifying Single-cell Multi-omics with Seamless Integration, Cross-modal Generation, and In-silico Exploration 97%
Similar papers in this journal
- Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST 97%
- uniPort: a unified computational framework for single-cell data integration with optimal transport 97%
- scMODAL: A general deep learning framework for comprehensive single-cell multi-omics data alignment with feature links 97%
Similar papers in this journal
- Spatial mutual nearest neighbors for spatial transcriptomics data 97%
- DESpace: spatially variable gene detection via differential expression testing of spatial clusters 96%
- ARTEMIS integrates autoencoders and schrodinger bridges to predict continuous dynamics of gene expression, cell population and perturbation from time-series single-cell data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.