Back

VCBench: A Multi-Dimensional Benchmark for Single-Cell Foundation Models

Weidener, L. S.; Brkic, M.; Jovanovic, M.; Ulgac, E.; Meduri, A.

2026-06-23 bioinformatics
10.64898/2026.06.18.733146 bioRxiv
Show abstract

Single-cell foundation models are increasingly positioned as virtual cells, yet their capabilities are assessed by fragmented, largely single-task benchmarks that obscure where these models improve on simple baselines. VCBench addresses this by synthesizing four independent virtual-cell frameworks into seven capability dimensions: perturbation response prediction, cross-species universality, gene regulatory network (GRN) inference, modality integration, temporal dynamics, multi-scale integration, and in silico experimentation. Each dimension is assessed for operational testability under current architectures and datasets: five admit direct or proxy evaluation, while multi-scale integration and in silico experimentation are structurally untestable as end-to-end tasks. We evaluate five foundation models (Geneformer, scGPT, UCE, TranscriptFormer, Arc State) against pre-registered linear and nearest-neighbor baselines across the five testable dimensions, and report three findings. First, the baselines match or exceed every foundation model on four of the five scored dimensions, replicating the reported competitiveness of linear baselines on perturbation prediction and extending it to cross-species transfer, GRN inference, and temporal ordering. Second, TranscriptFormer alone exceeds the strongest baseline on cross-modal RNA-to-protein prediction (53% Pearson improvement, with a documented contamination caveat) and is the only model to reach Level 2 in the pre-registered Virtual Cell (VC) Level rubric; the architectural choice behind this advantage simultaneously causes a spectral collapse that destroys its temporal-ordering performance, a tradeoff invisible to single-task benchmarks. Third, no foundation model publishes a complete cell-level training manifest, leaving data contamination undetectable to users. Alongside the benchmark, VCBench releases a Contamination Reporting Schema and contributes two further methodological tools: a common-label-set protocol that controls for class-count confounds in cross-species transfer, and a spread-error correlation probe for epistemic calibration.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
11.8%
2
Genome Biology
637 papers in training set
Top 1%
9.7%
3
Nature Methods
385 papers in training set
Top 1%
7.8%
4
Bioinformatics Advances
203 papers in training set
Top 0.5%
7.2%
5
NAR Genomics and Bioinformatics
242 papers in training set
Top 0.5%
6.7%
6
Nucleic Acids Research
1281 papers in training set
Top 4%
5.4%
7
Nature Communications
5641 papers in training set
Top 30%
4.8%
50% of probability mass above
8
Nature Biotechnology
172 papers in training set
Top 0.8%
4.8%
9
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.0%
10
Cell Systems
201 papers in training set
Top 2%
3.2%
11
PLOS Computational Biology
1863 papers in training set
Top 12%
2.4%
12
Scientific Reports
3612 papers in training set
Top 51%
1.9%
13
PLOS ONE
5266 papers in training set
Top 47%
1.9%
14
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
1.7%
15
GigaScience
212 papers in training set
Top 2%
1.7%
16
Patterns
78 papers in training set
Top 1%
1.7%
17
Genome Research
468 papers in training set
Top 4%
1.3%
18
Nature Machine Intelligence
70 papers in training set
Top 2%
1.1%
19
Cell Reports Methods
165 papers in training set
Top 3%
1.1%
20
iScience
1154 papers in training set
Top 26%
1.1%
21
Nature
645 papers in training set
Top 9%
1.0%
22
eLife
5828 papers in training set
Top 61%
1.0%
23
Molecular Systems Biology
162 papers in training set
Top 3%
1.0%
24
Cell Genomics
172 papers in training set
Top 4%
0.8%
25
Nature Computational Science
55 papers in training set
Top 2%
0.8%
26
Frontiers in Genetics
230 papers in training set
Top 6%
0.8%
27
Journal of Chemical Information and Modeling
238 papers in training set
Top 3%
0.6%