Back

Tokenizing single-cell transcriptomes as a native language for large language models

Xiao, C.; Ding, Y.; Bian, H.; Chen, Y.; Wei, L.; Zhang, X.

2026-07-11 bioinformatics
10.1101/2025.10.22.684047 bioRxiv
Show abstract

Large language models (LLMs) can process diverse forms of information once they are represented as tokens in a shared sequence space. However, single-cell transcriptomes remain a foreign modality to LLMs because they are continuous, high-dimensional molecular profiles rather than discrete linguistic units. Here, we propose CellTok, a tokenized single-cell language modeling approach that converts transcriptomic profiles into compact cellular token sequences and incorporates them into the vocabulary of a pretrained LLM. By representing cells as native tokens, CellTok enables cellular measurements, textual instructions, biological context, and multi-cell populations to be jointly processed within the same autoregressive modeling framework. Across diverse tasks, CellTok enable LLMs to recognize individual cells, interpret homogeneous and heterogeneous cell populations, infer disease-associated cellular states, predict cell-cell communication, model developmental trajectories, and generate cellular states. Moreover, prompt-based experiments show that providing appropriate biological context improves performance, indicating that CellTok can leverage LLM knowledge and contextual reasoning to support cellular data interpretation. These results demonstrate that single-cell transcriptomes can be transformed from a foreign molecular modality into a native language for LLMs, establishing a unified interface for modeling cells, populations, and biological knowledge in a shared token space.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Cell Systems
201 papers in training set
Top 0.1%
15.1%
2
Nature Communications
5641 papers in training set
Top 11%
15.1%
3
Nature Machine Intelligence
70 papers in training set
Top 0.1%
11.9%
4
Nature Methods
385 papers in training set
Top 1.0%
9.7%
50% of probability mass above
5
PLOS Computational Biology
1863 papers in training set
Top 9%
4.1%
6
Genome Biology
637 papers in training set
Top 4%
3.2%
7
Patterns
78 papers in training set
Top 0.7%
2.8%
8
Bioinformatics
1204 papers in training set
Top 6%
2.7%
9
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.4%
10
PLOS ONE
5266 papers in training set
Top 46%
1.9%
11
Nucleic Acids Research
1281 papers in training set
Top 8%
1.9%
12
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 27%
1.7%
13
Genome Research
468 papers in training set
Top 3%
1.7%
14
Molecular Systems Biology
162 papers in training set
Top 1%
1.7%
15
Scientific Reports
3612 papers in training set
Top 58%
1.5%
16
Nature Biotechnology
172 papers in training set
Top 3%
1.4%
17
iScience
1154 papers in training set
Top 25%
1.1%
18
BioData Mining
22 papers in training set
Top 0.6%
1.0%
19
eLife
5828 papers in training set
Top 60%
1.0%
20
BMC Bioinformatics
457 papers in training set
Top 5%
0.9%
21
npj Systems Biology and Applications
125 papers in training set
Top 2%
0.8%
22
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.8%
23
Genome Medicine
183 papers in training set
Top 5%
0.8%
24
Advanced Science
286 papers in training set
Top 11%
0.6%
25
Cell Reports Methods
165 papers in training set
Top 4%
0.6%
26
Genomics, Proteomics & Bioinformatics
16 papers in training set
Top 0.2%
0.6%
27
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.6%
28
Cell
431 papers in training set
Top 11%
0.6%