Proximal Exploration for Model-guided Protein Sequence Design
Ren, Z.; Li, J.; Ding, F.; Zhou, Y.; Ma, J.; Peng, J.
Show abstract
Designing protein sequences with a particular biological function is a long-lasting challenge for protein engineering. Recent advances in machine-learning-guided approaches focus on building a surrogate sequence-function model to reduce the burden of expensive in-lab experiments. In this paper, we study the exploration mechanism of model-guided sequence design. We leverage a natural property of protein fitness landscape that a concise set of mutations upon the wild-type sequence are usually sufficient to enhance the desired function. By utilizing this property, we propose Proximal Exploration (PEX) algorithm that prioritizes the evolutionary search for high-fitness mutants with low mutation counts. In addition, we develop a specialized model architecture, called Mutation Factorization Network (MuFacNet), to predict low-order mutational effects, which further improves the sample efficiency of model-guided evolution. In experiments, we extensively evaluate our method on a suite of in-silico protein sequence design tasks and demonstrate substantial improvement over baseline algorithms.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 94%
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 94%
- BatchDTA: Implicit batch alignment enhances deep learning-based drug-target affinity estimation 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 94%
- Building explainable graph neural network by sparse learning for the drug-protein binding prediction 94%
- Critiquing Protein Family Classification Models Using Sufficient Input Subsets 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.