Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization
Wan, L.; Bagherpoor Helabad, M.; Fraedrich, L.; Fleishman, S. J.; Weissenborn, M.
Show abstract
FuncLib and high-throughput FuncLib (htFuncLib) generate diverse, functional protein libraries using a stability-centered design; however, this substrate-independent approach lacks target-specific functional constraints. We developed a machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round. The system was benchmarked using previously published four-position fitness landscapes of three different proteins. The MLEE workflow successfully generated compact libraries enriched in globally high-fitness variants. After the initial training phase, an MLEE-enriched library of just 12 variants increased the hit rate for the global top-0.05% variants by 5- to 12-fold relative to the htFuncLib baseline. Screening a larger set of 96 variants recovered at least one of these top-performing enzymes in 61.3-99.4% of the simulations. We then applied MLEE to MthUPO-catalyzed {beta}-damascone hydroxylation. Across two rounds, 506 distinct variants were screened and sequenced. While the initial substrate-independent htFuncLib library yielded 14% of variants with activity above the wild type, the MLEE-enriched library increased this hit rate to 90% (97 of 108 variants) with activity above the wild type. The best variant increased the turnover number for 4-hydroxy-{beta}-damascone by 11.8-fold and achieved >99% regioisomeric excess. MLEE may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants. TABLE OF CONTENT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=100 SRC="FIGDIR/small/739108v1_ufig1.gif" ALT="Figure 1"> View larger version (18K): org.highwire.dtl.DTLVardef@b1f79dorg.highwire.dtl.DTLVardef@1f77429org.highwire.dtl.DTLVardef@eb4b6dorg.highwire.dtl.DTLVardef@1a4e5ae_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.