Back

Ultrahigh throughput screening to train generative protein models for engineering specificity into unspecific peroxygenases

Nair, P. M.; Steinberg, D. M.; Resende, T.; Suarez, A. F.; Power, H.; Ong, C. S.; Sairam, V.; Cheng, S. H.-Y.; Fragata, L.; Serafini, A.; Oh, V.; Tan, S. H.; Speight, R. E.; Vahidi, A. K.

2025-11-04 bioengineering
10.1101/2025.11.02.685536 bioRxiv
Show abstract

Enzyme engineering is central to developing biocatalysts with improved activity and specificity, yet traditional approaches are often limited by the scale of experimental screening. Here we combine ultrahigh-throughput microfluidic droplet screening with generative machine learning to engineer specificity into an unspecific peroxygenase (UPO) from Aspergillus brasiliensis (AbrUPO). A library of 5e6 variants was expressed in Komagataella phaffii (Pichia pastoris) and screened for increased production of styrene oxide and reduced production of phenylacetaldehyde and phenylpropanol. From an ultrahigh throughput screening dataset of >30,000 unique sequences paired with function data. These data were used both for direct selection of enriched variants and to train a task-specific generative long short-term memory (LSTM) model using the Variational Search Distributions (VSD) framework. We compared variants selected by rank aggregation from the screening data (R series) with novel sequences generated by the refined generative model (G series). Experimental assays revealed that multiple G variants outperformed both wildtype AbrUPO and R variants, in terms of styrene oxide production and specificity. Correlation analyses further showed that the task-specific generative model outperformed protein models pretrained on large publicly available datasets in predicting experimental outcomes. Our results demonstrate that coupling ultrahigh throughput screening with generative protein models enables efficient discovery of improved enzyme variants beyond those identified by direct screening alone, providing a scalable strategy for task-specific enzyme engineering. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=105 SRC="FIGDIR/small/685536v2_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@1226725org.highwire.dtl.DTLVardef@1a1cd7eorg.highwire.dtl.DTLVardef@1ba1a64org.highwire.dtl.DTLVardef@11ad176_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.