Back

NeuroFold: A Multimodal Approach to Generating Novel Protein Variants in silico

Amani, K.; Fish, M.; Smith, M.; Castroverde, C. D. M.

2024-03-14 bioinformatics
10.1101/2024.03.12.584504 bioRxiv
Show abstract

The generation of high-performance enzyme variants with desired physicochemical and functional properties presents a formidable challenge in the field of protein engineering. Existing in silico design methods are limited by inadequate training data, insufficient diversity within datasets, and suboptimal sampling techniques. Here, we introduce a novel approach that addresses these limitations and significantly improves the efficiency of generating functional enzyme variants. Using a multimodal approach, NeuroFold can leverage sequence, structural, and homology data during both sampling and discrimination phases, thereby enabling more diverse and informed sampling of the sequence space. Our model demonstrated a 40-fold increase in Spearman rank correlation as compared to large language models (LLMs) such as ESM-1v and empowers the rapid creation of high-quality enzyme variants, such as the {beta}-lactamase variants generated by NeuroFold in this study, which demonstrated increased thermostability and varying levels of activity. This pipeline represents a promising advancement in the field of enzyme engineering, offering a valuable tool for the development of novel enzymes with enhanced performance and desired chemical properties. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=101 SRC="FIGDIR/small/584504v1_ufig1.gif" ALT="Figure 1"> View larger version (21K): org.highwire.dtl.DTLVardef@6d8137org.highwire.dtl.DTLVardef@13e6c28org.highwire.dtl.DTLVardef@12ebff6org.highwire.dtl.DTLVardef@3ce466_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.