Back

NicheAgent: LLM-Guided Zero-Shot Niche Identification for Spatial Transcriptomics

Dip, S. A.; Zhang, L.

2025-12-12 bioinformatics
10.64898/2025.12.09.693287 bioRxiv
Show abstract

Spatial transcriptomics provides high-resolution maps of gene expression within intact tissue architecture, enabling the study of cellular niches, functional layers, and microenvironmental structure. Yet, accurately assigning niche or layer identities remains challenging across platforms such as 10x Visium, MERFISH, and STARmap due to batch variability, incomplete marker panels, and the lack of universally consistent domain boundaries. Existing methods including SpaGCN, BayesSpace, STAGATE, and DeepST rely heavily on supervised labels, dataset-specific fine-tuning, or deep representation learning, which often over-smooth boundaries, fail to generalize across technologies, and provide limited interpretability. We introduce NicheAgent, a zero-shot, training-free framework for spatial niche identification guided by lightweight large language models (LLMs). NicheAgent constructs interpretable nichecards for each region, encoding canonical marker genes and prototype expression centroids. Each cell is first assigned using a deterministic nearest-prototype rule based solely on gene expression and 2-hop spatial neighborhoods. Low-confidence assignments are then selectively reviewed and corrected by an LLM using only interpretable signals: marker-gene coherence, neighborhood label consistency, and an allowed label set. A final spatial smoothing step enforces local structural coherence. Applied across Visium, MERFISH, and STARmap tissues without any retraining or domain-specific supervision, NicheAgent achieves robust and biologically meaningful niche delineation, outperforming many supervised and graph-based baselines on homogeneity, completeness, and mutual information, while offering transparent reasoning traces for each corrected decision. Our results demonstrate that LLM-guided refinement, when coupled with lightweight rule-based prototypes, provides a scalable, explainable, and cross-platform alternative to heavy deep learning models for spatial transcriptomics annotation.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.