Back

Computational Design of Metallohydrolases

Kim, D.; Woodbury, S. M.; Ahern, W.; Tischer, D.; Hanikel, N.; Salike, S.; Yim, J.; Pellock, S. J.; Lauko, A.; Kalvet, I.; Hilvert, D.; Baker, D.

2025-04-11 biochemistry
10.1101/2024.11.13.623507 bioRxiv
Show abstract

De novo enzyme design starts from a description of an ideal active site composed of catalytic residues surrounding the reaction transition state(s), and builds a protein structure that contains this site1-7. Generative AI methods such as RFdiffusion11,12 now enable the direct generation of proteins around active sites, but to date, such scaffolding has required specification of both the position in the sequence and the backbone coordinates of each catalytic residue, which complicates sampling. Here we introduce a generative AI method called RFdiffusion2 that overcomes these limitations and use it to design zinc metallohydrolases starting from a density functional theory description of the active site geometry. Of an initial set of 96 designs tested experimentally, the most active has a kcat/KM of 16,000 M-1 s-1, orders of magnitude higher than previously designed metallohydrolases.6,7,13,14 A second round of 96 designs yielded 3 additional highly active enzymes, with kcat/KM up to 53,000 M-1 s-1 and kcat up to 1.5 s-1. The structures of the four enzymes are very different from each other and from the structures in the PDB. Each enzyme positions the reaction substrate almost perfectly for nucleophilic attack by a water molecule activated by the bound metal, and are predicted by PLACER15 and Chai-144 to have highly preorganized active sites. The ability to generate highly active catalysts straight out of the computer, without experimental optimization, using quantum chemistry calculated active site geometries should open the door to a new generation of potent designer enzymes.16,17

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.