Back

NanoCellAnnotator: Formalizing Expert Cell Type Annotation with Large Language Models

Mahmud, M. I.; Kochat, V.; Anzum, H.; Satpati, S.; Dwarampudi, J. M. R.; Rai, K.; Banerjee, T.

2026-06-25 bioinformatics
10.64898/2026.06.21.728965 bioRxiv
Show abstract

Motivation: Cell-type annotation in spatial transcriptomics is challenging due to sparse gene panels, spatial heterogeneity, and limited availability of tissue-matched reference atlases. Recent approaches have explored large language models (LLMs) for integrating biological knowledge during annotation, but unconstrained inference can produce biologically unsupported predictions and hallucinated cell types. In addition, many LLM-based pipelines rely on large cloud-hosted models that limit reproducibility and deployment in privacy-sensitive environments. Results: We introduce NanoCellAnnotator, a biologically constrained and confidence-aware framework for automated cell-type annotation in spatial transcriptomics. The framework de-couples spatial structure discovery, deterministic biological evidence construction, and language model-based semantic inference. Spatial clusters are identified using hybrid spatially regularized non-negative matrix factorization (hSNMF), after which cluster-level marker genes are abstracted into ontology-derived functional programs using Gene Ontology enrichment and GO-slim projection. A lightweight locally executable language model performs constrained label selection within a curated admissible label space derived from PanglaoDB and CellMarker. Annotation confidence is estimated independently using marker support strength and lineage separation, enabling ambiguous or heterogeneous clusters to be explicitly flagged. We evaluate NanoCellAnnotator on Xenium spatial transcriptomics data from intrahepatic cholangiocarci-noma and an independent breast cancer spatial transcriptomics dataset. The framework recovers canonical cell populations with high confidence while identifying heterogeneous or transitional spatial domains as ambiguous. Agreement with manual annotations was evaluated using accuracy and adjusted Rand index. Availability: Code available at https://github.com/ishtyaqmahmud/NanoCellAnnotator.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 1%
18.5%
2
Nature Biotechnology
172 papers in training set
Top 0.1%
11.9%
3
Nature Methods
385 papers in training set
Top 0.9%
9.8%
4
Genome Biology
637 papers in training set
Top 1%
7.9%
5
Nature Communications
5641 papers in training set
Top 22%
7.3%
50% of probability mass above
6
Genome Research
468 papers in training set
Top 1%
4.8%
7
Cell Systems
201 papers in training set
Top 1%
3.4%
8
PLOS Computational Biology
1863 papers in training set
Top 9%
3.4%
9
Nucleic Acids Research
1281 papers in training set
Top 6%
3.2%
10
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.2%
11
Cell Reports Methods
165 papers in training set
Top 0.9%
2.6%
12
Genome Medicine
183 papers in training set
Top 3%
1.7%
13
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 31%
1.4%
14
Advanced Science
286 papers in training set
Top 6%
1.3%
15
Patterns
78 papers in training set
Top 2%
1.3%
16
Bioinformatics Advances
203 papers in training set
Top 4%
1.1%
17
iScience
1154 papers in training set
Top 26%
1.1%
18
PLOS ONE
5266 papers in training set
Top 59%
1.0%
19
GigaScience
212 papers in training set
Top 4%
1.0%
20
Nature Genetics
286 papers in training set
Top 5%
0.8%
21
Molecular Systems Biology
162 papers in training set
Top 3%
0.8%
22
Scientific Reports
3612 papers in training set
Top 74%
0.8%
23
Communications Biology
993 papers in training set
Top 30%
0.8%
24
BMC Methods
15 papers in training set
Top 0.3%
0.6%
25
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.6%