Back

OCellus: A Language-Model Framework for Single-Cell, Spatial, and Perturbation Biology with Natural-Language Reasoning

Zhang, C.; Sun, J.; Xu, Z.; Liao, R.; Yin, A.; Gao, H.; Liu, E.; Bao, Y.; Zhao, L.; Wang, G.

2026-07-12 bioinformatics
10.64898/2026.07.08.737248 bioRxiv
Show abstract

Computational modeling of cellular behavior--the virtual cell--has emerged as a stated grand challenge at the intersection of artificial intelligence and biology, yet existing foundation models remain specialized: single-cell models process dissociated transcriptomes only, spatial models require dedicated spatial-aware architectures, and perturbation predictors depend on manually curated knowledge bases that cap generalization. Here we introduce OCellus, a single nine-billion-parameter language model (Qwen3.5-9B) fine-tuned on twenty-two biological tasks that simultaneously addresses all three limitations through three coordinated technical contributions on a shared backbone. First, EvenClock encodes two-dimensional spatial coordinates as eighteen clockface sectors of text, enabling spatial reasoning on a vanilla language model without architectural modification; on ten spatial transcriptomics tasks OCellus attains 77 percent spatial-neighborhood accuracy, 96 percent spatial-cellchat accuracy, and 0.70 proportion-cosine similarity on spatial deconvolution, all without any spatial-aware architectural components. Second, per-gene language-model embeddings replace the Gene Ontology annotations that GEARS depends on, achieving Pearson correlation 0.945 on the Replogle 2022 perturbation benchmark versus 0.84 for GEARS across 457 completely unseen knockout genes. Third, OCellus-Agent provides a Planner-Router-Verifier natural-language interface that achieves 75 percent pipeline accuracy on eighty multi-task queries. Removing language-model embeddings collapses perturbation Pearson to 0.06, confirming that learned functional representations--not graph topology--drive the gain. As a cell-type encoder, OCellus ranks first among fourteen foundation models in linear-probe accuracy at 95.1 percent across four benchmark datasets, and reaches 72.6 percent average across twenty-two evaluated biological tasks--a 57-percentage-point absolute gain over the strongest baseline configuration. As a language model, OCellus uniquely generates natural-language explanations of its predictions, a capability absent from all competing methods. Code, pre-trained model weights, the graph-neural-network module, and the agent system will be made available upon publication.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Nature Methods
385 papers in training set
Top 0.3%
18.9%
2
Nature
645 papers in training set
Top 1%
11.3%
3
Nature Biotechnology
172 papers in training set
Top 0.2%
9.9%
4
Nature Communications
5641 papers in training set
Top 19%
9.1%
5
Cell Systems
201 papers in training set
Top 0.3%
8.1%
50% of probability mass above
6
Science
477 papers in training set
Top 2%
5.0%
7
Bioinformatics
1204 papers in training set
Top 5%
3.6%
8
Genome Biology
637 papers in training set
Top 4%
2.4%
9
Nucleic Acids Research
1281 papers in training set
Top 8%
2.0%
10
PLOS Computational Biology
1863 papers in training set
Top 14%
1.8%
11
Cell
431 papers in training set
Top 6%
1.8%
12
Bioinformatics Advances
203 papers in training set
Top 3%
1.8%
13
Nature Genetics
286 papers in training set
Top 3%
1.7%
14
Genome Research
468 papers in training set
Top 4%
1.5%
15
PLOS ONE
5266 papers in training set
Top 54%
1.2%
16
Molecular Systems Biology
162 papers in training set
Top 2%
1.2%
17
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.0%
18
iScience
1154 papers in training set
Top 30%
1.0%
19
eLife
5828 papers in training set
Top 62%
0.9%
20
Nature Machine Intelligence
70 papers in training set
Top 2%
0.9%
21
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
0.9%
22
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 40%
0.9%
23
Cell Genomics
172 papers in training set
Top 4%
0.9%
24
Cell Reports
1498 papers in training set
Top 28%
0.6%
25
Science Advances
1243 papers in training set
Top 32%
0.6%
26
Patterns
78 papers in training set
Top 3%
0.6%
27
ACS Synthetic Biology
287 papers in training set
Top 2%
0.6%