Back

Acquiring Improved Protein Variants With Probabilistic Preferential Learning

van der Flier, F. J.; de Ridder, D.; Probst, D.; Redestig, H.

2026-06-26 bioinformatics
10.64898/2026.06.22.733688 bioRxiv
Show abstract

Variant effect prediction (VEP) models can be used to select promising novel enzymes from a pool of candidates. Most supervised VEP models are framed as regression tasks, placing more emphasis on getting the predicted quantities correct than on the relative comparison of individual candidates. Preferential or contrastive models may better align with the goal of selection, or acquisition, especially when informed by predictive uncertainty. Here, we introduce a probabilistic preferential learning model based on the Kermut Gaussian process (PKermut) that we designed with the ambition to increase the hit rate among selected variants. We benchmark PKermut against established models, including the original Kermut, the RITA regressor, and an augmented Potts model, on 69 curated ProteinGym datasets across various assay categories. To evaluate acquisition performance, we propose a novel quantile cross-validation scheme that ensures the evaluation of a models ability to extrapolate by reserving high-performing variants exclusively for the test set. We assess models using Spearman correlation and evaluate their acquisition performance using five different acquisition functions, encompassing both uncertainty-aware and unaware strategies. Our experimental results indicate that uncertainty estimates improve the acquisition ability of our models, and that strategies that reward uncertainty generally result in better outcomes than those that do not on single-mutation variant datasets. We observe that PKermuts Spearman scores and ability to acquire improved variants are greatly affected by the number of variant comparisons sampled in the training set. Kermut achieves the highest Spearman correlation in 54/69 datasets (78%), compared to 12/69 (17%) for PKermut. For acquisition performance, Kermut leads in 44/69 datasets (64%), while PKermut leads in 15/69 (22%). While at this stage PKermut is not a recommended alternative to Kermut, its contrastive nature offers several conceptual opportunities. We share our findings to inspire further development aimed at improving the alignment between training objectives of VEP models and their downstream application in protein engineering.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 1%
19.1%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.4%
13.3%
3
Bioinformatics Advances
203 papers in training set
Top 0.3%
8.1%
4
Bioinformatics
1204 papers in training set
Top 3%
7.0%
5
Briefings in Bioinformatics
354 papers in training set
Top 1%
7.0%
50% of probability mass above
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 1%
3.4%
7
Journal of Cheminformatics
29 papers in training set
Top 0.2%
3.3%
8
Scientific Reports
3612 papers in training set
Top 41%
2.5%
9
Nature Communications
5641 papers in training set
Top 39%
2.5%
10
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.6%
2.5%
11
BMC Bioinformatics
457 papers in training set
Top 3%
2.2%
12
PLOS ONE
5266 papers in training set
Top 48%
1.7%
13
International Journal of Molecular Sciences
494 papers in training set
Top 7%
1.7%
14
BMC Genomics
406 papers in training set
Top 5%
1.5%
15
Journal of Molecular Biology
232 papers in training set
Top 2%
1.5%
16
Protein Science
246 papers in training set
Top 3%
1.2%
17
Molecular Systems Biology
162 papers in training set
Top 2%
1.2%
18
Nature Machine Intelligence
70 papers in training set
Top 2%
1.2%
19
Cell Systems
201 papers in training set
Top 4%
1.0%
20
Communications Chemistry
48 papers in training set
Top 1%
1.0%
21
Frontiers in Bioinformatics
49 papers in training set
Top 1%
0.9%
22
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
0.9%
23
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 40%
0.9%
24
Biophysical Journal
631 papers in training set
Top 4%
0.9%
25
Nucleic Acids Research
1281 papers in training set
Top 14%
0.6%
26
Chemical Science
73 papers in training set
Top 2%
0.6%
27
Cell Reports Methods
165 papers in training set
Top 4%
0.6%
28
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.4%
0.5%
29
BioData Mining
22 papers in training set
Top 1%
0.5%
30
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 2%
0.5%