Back

AI-guided discovery for low-resource peptide engineering using evolutionary scale modeling

Andrekson, L.; Rydbergh, R.; Mercado, R.; Wenzel, M.

2026-07-01 bioinformatics
10.64898/2026.06.25.734678 bioRxiv
Show abstract

Reliable estimation of downstream performance in low-data peptide machine learning is critical for guiding early-stage AI-driven peptide engineering. Yet, it is often unclear how to assess whether a model will be effective in iterative discovery settings. Here, we show that the cross validation R2 score can serve as a simple and robust proxy for predicting active learning workflow performance, enabling early-stage evaluation of model suitability for sequential peptide optimization. To support this, we introduce SCARSE, a machine learning framework combining ESM-2 protein language model embeddings with Gaussian process regression and extremely randomized trees classification, designed for low-resource peptide property prediction (20-500 training samples). We benchmark SCARSE across 23 peptide and small-protein datasets covering substitution and indel variants, antimicrobial peptides, cell-penetrating peptides, and toxic/non-toxic peptides. SCARSE significantly outperforms a hand-engineered descriptor baseline on substitution and indel tasks, while comparable performance was achieved on shorter peptide non-mutant datasets where simpler descriptors capture enough of the signal. In simulated active learning workflows, SCARSE consistently outperforms baseline and random sampling strategies. Notably, we demonstrate that CV R2 computed from as few as 50 labeled peptides can be sufficient to estimate final active learning end-point performance, providing a practical, data-efficient criterion for deciding whether a given dataset combined with SCARSE is suitable for iterative peptide discovery. SCARSE is released as a pip package and is available via HuggingFace Spaces to facilitate integration into peptide engineering workflows.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.2%
22.4%
2
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.2%
9.8%
3
Briefings in Bioinformatics
354 papers in training set
Top 0.8%
7.9%
4
Nature Machine Intelligence
70 papers in training set
Top 0.2%
7.9%
5
PLOS Computational Biology
1863 papers in training set
Top 7%
4.9%
50% of probability mass above
6
mAbs
32 papers in training set
Top 0.1%
4.9%
7
Protein Science
246 papers in training set
Top 0.9%
4.3%
8
Communications Chemistry
48 papers in training set
Top 0.1%
4.1%
9
Nature Communications
5641 papers in training set
Top 35%
3.2%
10
Chemical Science
73 papers in training set
Top 0.5%
2.8%
11
Bioinformatics
1204 papers in training set
Top 7%
1.7%
12
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
1.7%
13
Cell Systems
201 papers in training set
Top 3%
1.4%
14
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 32%
1.3%
15
Cell Reports Methods
165 papers in training set
Top 3%
1.1%
16
Nature Methods
385 papers in training set
Top 5%
1.1%
17
Journal of Molecular Biology
232 papers in training set
Top 3%
1.1%
18
Scientific Reports
3612 papers in training set
Top 65%
1.1%
19
Bioinformatics Advances
203 papers in training set
Top 4%
1.1%
20
PLOS ONE
5266 papers in training set
Top 57%
1.1%
21
Nucleic Acids Research
1281 papers in training set
Top 12%
1.0%
22
Advanced Science
286 papers in training set
Top 8%
1.0%
23
Journal of Cheminformatics
29 papers in training set
Top 0.6%
1.0%
24
eLife
5828 papers in training set
Top 64%
0.8%
25
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 1%
0.6%