Back

Machine Learning enables efficient and effective affinity maturation of nanobodies

Paul, S. B.; Harvey, E. P.; Osei-Owusu, J.; Kollasch, A. W.; Riesselman, A. J.; McMahon, C.; Gazizov, A.; Anuganti, M.; Belay, F.; Kieu, M. A.; Zhu, H.; Hollingsworth, L. R.; Harper, J. W.; Moshinsky, D. J.; Teixeira, A. R. R.; Marks, D. S.; Kruse, A. C.

2026-01-12 bioinformatics
10.64898/2026.01.11.698911 bioRxiv
Show abstract

Antibodies can bind their targets with exquisite potency and selectivity due in part to large antibody-target protein-protein interaction surface areas. Despite the very large size and diversity of synthetic libraries, in vitro sorting alone tends to yield binders with modest affinities. By analogy to the in vivo affinity maturation in the natural immune system, these initial hits are typically affinity matured in vitro to achieve high affinity binding. However, affinity maturation campaigns can be laborious, often requiring multiple selection rounds and strategies for each clone to be optimized. Here, we investigated whether one could accelerate the discovery of optimized binders using machine learning on sequencing data from single selection sorts of affinity maturation yeast-display campaigns. Our results show that sparse sequencing data from a single sorting round can predict sequences that are enriched after multiple rounds. We also find that linear models outperform deep neural networks and semi-supervised approaches in ranking validated affinity-enhancing substitutions. Linear models are also more interpretable, offering insights into residue preferences that can be leveraged for further engineering. We use our models to design and select optimized nanobody binders to relaxin family peptide receptor 1 (RXFP1), yielding multiple improved binders including 3 sub nanomolar binders with the best exhibiting a [~]2500-fold improvement over WT.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

1
mAbs
32 papers in training set
Top 0.1%
38.5%
2
Nature Communications
5641 papers in training set
Top 14%
12.5%
50% of probability mass above
3
Cell Systems
201 papers in training set
Top 0.3%
8.7%
4
Nature Biotechnology
172 papers in training set
Top 1%
3.9%
5
Nature Methods
385 papers in training set
Top 3%
3.4%
6
Science
477 papers in training set
Top 3%
2.7%
7
eLife
5828 papers in training set
Top 38%
2.7%
8
Cell Reports Methods
165 papers in training set
Top 1%
2.1%
9
Protein Science
246 papers in training set
Top 2%
1.7%
10
PLOS Computational Biology
1863 papers in training set
Top 15%
1.7%
11
Nature Machine Intelligence
70 papers in training set
Top 2%
1.6%
12
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 31%
1.5%
13
Science Advances
1243 papers in training set
Top 23%
1.4%
14
Cell Genomics
172 papers in training set
Top 3%
1.1%
15
Molecular Systems Biology
162 papers in training set
Top 2%
1.0%
16
iScience
1154 papers in training set
Top 30%
1.0%
17
Bioinformatics
1204 papers in training set
Top 8%
1.0%
18
Scientific Reports
3612 papers in training set
Top 72%
0.9%
19
Communications Biology
993 papers in training set
Top 28%
0.9%
20
Nucleic Acids Research
1281 papers in training set
Top 13%
0.9%
21
Patterns
78 papers in training set
Top 2%
0.9%
22
Cell Chemical Biology
94 papers in training set
Top 2%
0.8%
23
Cell Reports
1498 papers in training set
Top 28%
0.8%
24
Structure
193 papers in training set
Top 3%
0.6%
25
Nature Chemical Biology
119 papers in training set
Top 3%
0.6%