Back

Overestimating zero-shot fitness prediction: Broad benchmarks mask local failures and practical limitations

Woolley, P. R.; Feller, A.; Ellington, A. O.; Wilke, C. O.

2026-06-07 bioengineering
10.64898/2026.06.04.730121 bioRxiv
Show abstract

Deep learning models have emerged as promising tools for navigating mutational landscapes in protein engineering. These models can be used to predict mutation fitness without the need for task-specific training, a process known as zero-shot prediction. However, their practical utility remains only partially characterized. Here, we evaluate the zero-shot performance of a panel of protein sequence and structure models across a range of benchmarking conditions, focusing on factors that complicate the interpretation of aggregate metrics. We show that input modality (sequence vs. structure) does not dictate performance on phenotypic tasks. Instead, performance is sensitive to experimental variability and is heavily confounded by correlation between phenotype and protein abundance. While available models may act as coarse filters separating fit mutations from deleterious ones, they cannot meaningfully rank a set of fit mutations or prioritize new-to-nature functions. Ultimately, the practical utility of zero-shot prediction from protein models is narrower than aggregate benchmarks imply.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 1%
18.4%
2
Cell Systems
201 papers in training set
Top 0.2%
12.4%
3
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 3%
9.7%
4
Nature Communications
5641 papers in training set
Top 24%
6.7%
5
eLife
5828 papers in training set
Top 21%
5.5%
50% of probability mass above
6
Scientific Reports
3612 papers in training set
Top 17%
5.5%
7
Journal of Chemical Information and Modeling
238 papers in training set
Top 1.0%
5.1%
8
Nature Machine Intelligence
70 papers in training set
Top 0.8%
3.4%
9
Protein Engineering, Design and Selection
15 papers in training set
Top 0.1%
3.2%
10
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
3.2%
11
Protein Science
246 papers in training set
Top 2%
2.4%
12
Briefings in Bioinformatics
354 papers in training set
Top 4%
1.7%
13
PLOS ONE
5266 papers in training set
Top 48%
1.7%
14
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.9%
1.3%
15
International Journal of Molecular Sciences
494 papers in training set
Top 11%
1.1%
16
ACS Synthetic Biology
287 papers in training set
Top 2%
1.1%
17
Biophysical Journal
631 papers in training set
Top 4%
1.1%
18
Journal of Cheminformatics
29 papers in training set
Top 0.6%
1.0%
19
Nature Methods
385 papers in training set
Top 6%
1.0%
20
Science Advances
1243 papers in training set
Top 28%
1.0%
21
PRX Life
42 papers in training set
Top 1%
0.6%