Back

Protein CREATE enables closed-loop design of de novo synthetic protein binders

Lourenco, A. L.; Subramanian, A. M.; Spencer, R. K.; Miao, J.; Anaya, M.; Fu, W.; Chow, E. D.; Thomson, M.

2024-12-22 bioengineering
10.1101/2024.12.20.629847 bioRxiv
Show abstract

Proteins have proven to be useful agents in a variety of fields, from serving as potent therapeutics to enabling complex catalysis for chemical manufacture. However, they remain difficult to design and are instead typically selected for using extensive screens or directed evolution. Recent developments in protein large language models have enabled fast generation of diverse protein sequences in unexplored regions of protein space predicted to fold into varied structures, bind relevant targets, and catalyze novel reactions. Nevertheless, we lack methods to characterize these proteins experimentally at scale and update generative models based on those results. We describe Protein CREATE (Computational Redesign via an Experiment-Augmented Training Engine), an integrated computational and experimental pipeline that incorporates an experimental workflow leveraging next generation sequencing and phage display with single-molecule readouts to collect vast amounts of quantitative binding data for updating protein large language models. We use Protein CREATE to generate and assay thousands of designed binders to IL-7 receptor and insulin receptor with parallel positive and negative selections to identify on-target binders. We discover not only individual novel binders but also features of ligand-receptor binding, including preservation of the IL7R - ligand hydrophobic interface specifically and existence of multiple approaches to contact the insulin receptor. We also demonstrate the importance of structural features, such as the lack of unpaired cysteine residues, toward design fidelity and find computational pre-screening metrics, such as interchain predicted TM scoring (iPTM), while useful, are imperfect predictors as they neither guarantee experimental binding nor rule it out. We use the data collected from Protein CREATE to score designs from the initial generative models. Globally, Protein CREATE will power future closed-loop design-build-test cycles to enable fine-grained design of protein binders.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Protein Science
246 papers in training set
Top 0.1%
18.8%
2
Protein Engineering, Design and Selection
15 papers in training set
Top 0.1%
13.1%
3
PLOS Computational Biology
1863 papers in training set
Top 6%
6.3%
4
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 12%
4.4%
5
Nature Methods
385 papers in training set
Top 2%
4.1%
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 1.0%
4.1%
50% of probability mass above
7
Cell Systems
201 papers in training set
Top 1%
4.1%
8
Nature Communications
5641 papers in training set
Top 35%
3.3%
9
Structure
193 papers in training set
Top 0.9%
2.7%
10
mAbs
32 papers in training set
Top 0.2%
2.4%
11
Angewandte Chemie International Edition
93 papers in training set
Top 0.7%
2.4%
12
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.6%
1.9%
13
Nature Computational Science
55 papers in training set
Top 0.6%
1.7%
14
Journal of Molecular Biology
232 papers in training set
Top 2%
1.7%
15
ACS Chemical Biology
167 papers in training set
Top 1%
1.7%
16
Biochemical Journal
91 papers in training set
Top 0.9%
1.4%
17
Science Advances
1243 papers in training set
Top 24%
1.1%
18
eLife
5828 papers in training set
Top 56%
1.1%
19
Advanced Science
286 papers in training set
Top 6%
1.1%
20
Scientific Reports
3612 papers in training set
Top 64%
1.1%
21
Journal of the American Chemical Society
217 papers in training set
Top 2%
1.1%
22
PLOS ONE
5266 papers in training set
Top 54%
1.1%
23
ACS Synthetic Biology
287 papers in training set
Top 2%
1.1%
24
Cell Chemical Biology
94 papers in training set
Top 1%
1.1%
25
Bioinformatics
1204 papers in training set
Top 9%
0.9%
26
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
0.9%
27
Nature Biotechnology
172 papers in training set
Top 4%
0.9%
28
iScience
1154 papers in training set
Top 34%
0.9%
29
Cell Genomics
172 papers in training set
Top 4%
0.6%
30
Science
477 papers in training set
Top 10%
0.6%