Back

CatESO: Differentiable Enzyme Sequence Optimization Guided by Substrate-Aware kcat Prediction

Gan, Z.; Xu, Y.; Xu, J.; Wu, Z.; Huang, J.; Yin, J.; Chen, G.; Zhang, J. Z. H.

2026-07-06 biochemistry
10.64898/2026.07.04.736506 bioRxiv
Show abstract

Enzymes drive biological chemistry and offer greener routes to chemicals, materials and medicines, yet their broader use as biocatalysts is often limited by insufficient catalytic turnover. Improving turnover is hard: measured rate constants are scarce and protein sequence space is vast. Deep-learning models now predict the turnover number, Kcat, with growing accuracy, but they are typically applied after sequence generation to score or filter candidates, which separates the kinetic objective from the design itself. To bridge the gap between sequence generation and kinetic evaluation, we introduce CatESO, a differentiable sequence optimizer that enables direct, gradient-guided design of substrate-specific catalytic turnover. By backpropagating through a cross-modal Kcat predictor under continuous sequence relaxation, CatESO co-optimizes predicted catalytic activity, evolutionary plausibility and structural integrity in one end-to-end framework, using ESM-2 and ESMFold to keep designs evolutionarily plausible and foldable. Across seven stringent out-of-distribution enzymes spanning EC classes 1-7, CatESO raised model-predicted Kcat for the vast majority of designs, with a median predicted fold change of 1.52 while every variant retained a pLDDT above 70. Against RFdiffusion3-LigandMPNN pipeline and ZymCtrl, CatESO struck a better balance between predicted activity and structural confidence. By making substrate-conditioned kinetic objectives differentiable, CatESO carries differentiable protein design beyond structure- and binding-centred goals to enzyme catalytic function, giving a general route to function-oriented enzyme engineering.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Nature
645 papers in training set
Top 0.8%
15.0%
2
Nature Communications
5641 papers in training set
Top 14%
12.8%
3
ACS Catalysis
18 papers in training set
Top 0.1%
9.6%
4
Science
477 papers in training set
Top 1%
5.5%
5
Nature Chemical Biology
119 papers in training set
Top 0.4%
5.5%
6
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
4.0%
50% of probability mass above
7
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 13%
4.0%
8
Nature Chemistry
42 papers in training set
Top 0.3%
3.2%
9
Communications Chemistry
48 papers in training set
Top 0.2%
3.2%
10
Nature Machine Intelligence
70 papers in training set
Top 1.0%
2.8%
11
Angewandte Chemie International Edition
93 papers in training set
Top 0.7%
2.6%
12
Chemical Science
73 papers in training set
Top 0.7%
2.4%
13
Cell Systems
201 papers in training set
Top 2%
2.4%
14
ACS Central Science
71 papers in training set
Top 0.5%
2.4%
15
Cell Chemical Biology
94 papers in training set
Top 0.9%
1.7%
16
eLife
5828 papers in training set
Top 52%
1.5%
17
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
18
Nature Biotechnology
172 papers in training set
Top 4%
1.1%
19
Nature Methods
385 papers in training set
Top 6%
1.0%
20
Protein Science
246 papers in training set
Top 3%
1.0%
21
mAbs
32 papers in training set
Top 0.5%
0.8%
22
ACS Chemical Biology
167 papers in training set
Top 2%
0.8%
23
ChemMedChem
16 papers in training set
Top 0.4%
0.6%
24
Cell Reports Methods
165 papers in training set
Top 5%
0.6%
25
Journal of the American Chemical Society
217 papers in training set
Top 3%
0.6%
26
Nature Biomedical Engineering
47 papers in training set
Top 2%
0.6%
27
npj Digital Medicine
118 papers in training set
Top 4%
0.6%