Back

Context-Dependent Design of Induced-fit Enzymes using Deep Learning Generates Well Expressed, Thermally Stable and Active Enzymes

Zimmerman, L.; Alon, N.; Levin, I.; Koganitsky, A.; Brestel, C.; Lapidoth, G. D.

2023-07-29 bioinformatics
10.1101/2023.07.27.550799 bioRxiv
Show abstract

The potential of engineered enzymes in practical applications is often constrained by limitations in their expression levels, thermal stability, and the diversity and magnitude of catalytic activities. De-novo enzyme design, though exciting, is challenged by the complex nature of enzymatic catalysis. An alternative promising approach involves expanding the capabilities of existing natural enzymes to enable functionality across new substrates and operational parameters. To this end we introduce CoSaNN (Conformation Sampling using Neural Network), a novel strategy for enzyme design that utilizes advances in deep learning for structure prediction and sequence optimization. By controlling enzyme conformations, we can expand the chemical space beyond the reach of simple mutagenesis. CoSaNN uses a context-dependent approach that accurately generates novel enzyme designs by considering non-linear relationships in both sequence and structure space. Additionally, we have further developed SolvIT, a graph neural network trained to predict protein solubility in E.Coli, as an additional optimization layer for producing highly expressed enzymes. Through this approach, we have engineered novel enzymes exhibiting superior expression levels, with 54% of our designs expressed in E.Coli, and increased thermal stability with more than 30% of our designs having a higher Tm than the template enzyme. Furthermore, our research underscores the transformative potential of AI in protein design, adeptly capturing high order interactions and preserving allosteric mechanisms in extensively modified enzymes. These advancements pave the way for the creation of diverse, functional, and robust enzymes, thereby opening new avenues for targeted biotechnological applications.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Protein Science
246 papers in training set
Top 0.1%
40.6%
2
Nature Communications
5641 papers in training set
Top 23%
6.9%
3
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.1%
6.9%
50% of probability mass above
4
Scientific Reports
3612 papers in training set
Top 15%
5.6%
5
ACS Synthetic Biology
287 papers in training set
Top 0.7%
5.3%
6
Journal of Molecular Biology
232 papers in training set
Top 2%
2.0%
7
Biochemistry
148 papers in training set
Top 1%
1.8%
8
Nucleic Acids Research
1281 papers in training set
Top 10%
1.4%
9
The FEBS Journal
93 papers in training set
Top 1.0%
1.4%
10
ACS Omega
105 papers in training set
Top 2%
1.2%
11
The Journal of Physical Chemistry B
167 papers in training set
Top 1%
1.2%
12
Communications Biology
993 papers in training set
Top 20%
1.2%
13
Protein Engineering, Design and Selection
15 papers in training set
Top 0.1%
1.2%
14
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
15
Journal of Biological Chemistry
690 papers in training set
Top 8%
1.1%
16
Chemical Science
73 papers in training set
Top 1%
1.1%
17
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 2%
0.9%
18
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 40%
0.9%
19
Bioinformatics
1204 papers in training set
Top 8%
0.9%
20
Briefings in Bioinformatics
354 papers in training set
Top 7%
0.9%
21
PLOS ONE
5266 papers in training set
Top 60%
0.9%
22
Metabolic Engineering
75 papers in training set
Top 0.7%
0.6%
23
Structure
193 papers in training set
Top 2%
0.6%
24
ACS Catalysis
18 papers in training set
Top 0.2%
0.6%
25
Biophysical Journal
631 papers in training set
Top 5%
0.6%