Back

TissueNarrator: Generative Modeling of Spatial Transcriptomics with Large Language Models

Liu, S.; Tang, J.; Ma, J.; Liang, S.

2025-11-27 bioinformatics
10.1101/2025.11.24.690325 bioRxiv
Show abstract

The intricate spatial organization and molecular communication among cells are fundamental to multicellular systems. Spatial transcriptomics (ST) enables gene expression profiling while preserving spatial context, providing rich data for studying cellular interactions and tissue dynamics. However, most existing computational approaches focus on embedding-based tasks and provide limited generative capacity for simulating cell behavior in situ. Moreover, accurately interpreting spatial interactions requires extensive biological knowledge, which current models do not incorporate. Here, we introduce TO_SCPLOWISSUEC_SCPLOWNO_SCPLOWARRATORC_SCPLOW, a framework that reformulates spatial omics analysis as a language modeling problem. By representing tissue sections as spatial sentences - rank-based gene lists augmented with spatial coordinates and metadata - TO_SCPLOWISSUEC_SCPLOWNO_SCPLOWARRATORC_SCPLOW leverages pretrained large language models (LLMs) to learn spatially conditioned gene expression patterns. The model generates realistic, context-aware cellular profiles, predicts intercellular interactions, and performs in silico perturbation analyses. Across multiple ST technologies (MERFISH, Perturb-FISH, and CosMx SMI), TO_SCPLOWISSUEC_SCPLOWNO_SCPLOWARRATORC_SCPLOW achieves superior quantitative performance and recovers biologically meaning-ful ligand-receptor and signaling pathways. Furthermore, a conversational inference mode enables natural-language querying of tissue organization. By integrating pretrained biological knowledge with spatial context, TO_SCPLOWISSUEC_SCPLOWNO_SCPLOWARRATORC_SCPLOW establishes a new, scalable generative paradigm for modeling, simulating, and reasoning about tissue systems.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.