Back

Evaluating the ability of spatial transcriptomics foundation models to learn multi-scale spatial variation

Handa, D.; Martin-Linares, C.; Stein-O'Brien, G.; Ling, J.; Chitra, U.

2026-08-06 bioinformatics
10.64898/2026.08.01.742217 bioRxiv
Show abstract

Spatial gene expression results from the superposition of multiple sources of variation in gene expression across different spatial scales, including local microenvironment-associated variation and global spatial gradients. Spatial foundation models (SFMs) are large-scale machine learning models trained on cohorts of spatial transcriptomics (ST) data that, in principle, learn the different sources of spatial variation in gene expression. However, the embeddings learned by SFMs are difficult to interpret, and it remains unclear whether they fully capture such spatial variation. Here, we develop SAFFRON, a sparse autoencoder (SAE)-based framework for interpreting and evaluating SFMs. SAFFRON uses a Matryoshka SAE to decompose dense SFM embeddings into sparse, human-interpretable features and evaluates whether these features correlate with known sources of spatial variation. Using SAFFRON, we systematically benchmark the ability of several recent SFMs to identify local and global spatial variation in gene expression. We find that one SFM, Novae, learns global spatial gradients more accurately than naive, non-foundation model baselines, and that these gradients are concentrated in a small subset of sparse and human-interpretable SAE features revealed by SAFFRON. On the other hand, no SFM learns local microenvironment-associated patterns more accurately than such baselines. Our findings suggest that current SFMs do not systematically learn multi-scale spatial variation in gene expression. CodeSAFFRON is available at https://github.com/chitra-lab/SAFFRON.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.