Back

Machine Learning-Assisted Evolution of Broadly Functional Enzyme Libraries

Lal, R.; Yang, J.; Zhang, Z.; Arnold, F. H.

2026-07-24 bioengineering
10.64898/2026.07.23.740427 bioRxiv
Show abstract

Biocatalysis offers sustainable solutions to pressing challenges in chemical synthesis by exploiting the remarkable efficiency and selectivity of enzymes. Importantly, enzymes are able to accommodate non-native substrates and mediate transformations outside of their natural repertoire. Enzymes can be engineered for diverse applications by harnessing these promiscuous activities and optimizing them using directed evolution (DE). The success of a DE campaign, however, depends on the availability of a protein starting point that displays detectable levels of the desired function. To find a starting point, researchers often screen libraries of protein variants for novel activities, typically with low rates of success. Here, instead, we diversified the active site of a desirable parent protein and applied machine learning to generate informed, promiscuous libraries of protein variants. Specifically, we tested 26 different carbene and nitrene transfer reactions and used active learning-assisted directed evolution (ALDE) to generate optimized protoglobin variants with high activity across multiple reactions. We observed improvements in activity and selectivity for every reaction performed by the parent enzyme in at least one member of the ALDE-predicted libraries. Moreover, variants from these libraries can catalyze 5 out of 10 reactions not catalyzed by the parent protoglobin. These results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=99 SRC="FIGDIR/small/740427v1_ufig1.gif" ALT="Figure 1"> View larger version (31K): org.highwire.dtl.DTLVardef@6d5cfborg.highwire.dtl.DTLVardef@1f38e0aorg.highwire.dtl.DTLVardef@f25fa2org.highwire.dtl.DTLVardef@64a0dc_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.