Back

Tabular Foundation Models Are Competitive Cellular Perturbation Predictors Across Biological Scales

Palla, G.; Hillsley, A.; Kim, Y.-J.; Royer, L. A.

2026-07-01 bioinformatics
10.64898/2026.06.28.735106 bioRxiv
Show abstract

Predicting how cells respond to genetic and chemical perturbations is a central challenge in drug discovery and functional genomics. A growing ecosystem of specialized single-cell foundation models has been developed to address this problem, yet their practical advantage over domain-agnostic approaches remains unclear. Here we evaluate the power of Tabular Foundation Models such as TabICL and TabPFN, general-purpose pre-trained regression models, against domain-specific architectures including PRESAGE, scGPT, scLAMBDA, STACK and Prophet across four complementary evaluation settings: cell-level in-context cross-cell-type prediction, pseudobulk perturbation prediction on five Perturb-seq datasets of cell-lines, a genome-wide CRISPR screen in primary human CD4+ T cells, and embryo-level cell-type composition prediction in a zebrafish developmental perturbation atlas. In the cell-level cross-cell type perturbation prediction, Tabular Foundation Models perform on par or better than specialized models. On pseudobulk perturbation prediction, Tabular Foundation Models consistently outperform specialized baselines across multiple evaluation metrics and datasets. On whole-emrbryo cell-type composition prediction, Tabular Foundation Models are competitive with specialized baselines. These results demonstrate that general-purpose tabular in-context learning provides a strong and scalable alternative to bespoke biological architectures for perturbation response modeling across cell systems and scales.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 12%
14.7%
2
Cell Systems
201 papers in training set
Top 0.2%
11.6%
3
Molecular Systems Biology
162 papers in training set
Top 0.1%
9.5%
4
Briefings in Bioinformatics
354 papers in training set
Top 0.8%
8.7%
5
Genome Biology
637 papers in training set
Top 1%
7.7%
50% of probability mass above
6
Nature Machine Intelligence
70 papers in training set
Top 0.3%
7.7%
7
PLOS Computational Biology
1863 papers in training set
Top 11%
3.1%
8
Bioinformatics
1204 papers in training set
Top 5%
3.1%
9
Nature Methods
385 papers in training set
Top 3%
2.4%
10
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.3%
11
npj Systems Biology and Applications
125 papers in training set
Top 0.9%
2.1%
12
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
13
Bioinformatics Advances
203 papers in training set
Top 3%
1.7%
14
GigaScience
212 papers in training set
Top 3%
1.4%
15
Patterns
78 papers in training set
Top 2%
1.3%
16
Genome Research
468 papers in training set
Top 4%
1.3%
17
BMC Bioinformatics
457 papers in training set
Top 5%
1.1%
18
Genome Medicine
183 papers in training set
Top 4%
1.1%
19
Scientific Reports
3612 papers in training set
Top 70%
1.0%
20
eLife
5828 papers in training set
Top 61%
1.0%
21
iScience
1154 papers in training set
Top 32%
1.0%
22
Advanced Science
286 papers in training set
Top 9%
0.9%
23
Cell Genomics
172 papers in training set
Top 4%
0.8%
24
Nature Biotechnology
172 papers in training set
Top 4%
0.8%
25
BMC Genomics
406 papers in training set
Top 9%
0.8%
26
Cell Reports Methods
165 papers in training set
Top 4%
0.8%
27
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 46%
0.6%
28
Communications Biology
993 papers in training set
Top 37%
0.6%