Evolutionary context-integrated deep sequence modeling for protein engineering
Luo, Y.; Vo, L.; Ding, H.; Su, Y.; Liu, Y.; Qian, W.; Zhao, H.; Peng, J.
Show abstract
Protein engineering seeks to design proteins with improved or novel functions. Compared to rational design and directed evolution approaches, machine learning-guided approaches traverse the fitness landscape more effectively and hold the promise for accelerating engineering and reducing the experimental cost and effort. A critical challenge here is whether we are capable of predicting the function or fitness of unseen protein variants. By learning from the sequence and large-scale screening data of characterized variants, machine learning models predict functional fitness of sequences and prioritize new variants that are very likely to demonstrate enhanced functional properties, thereby guiding and accelerating rational design and directed evolution. While existing generative models and language models have been developed to predict the effects of mutation and assist protein engineering, the accuracy of these models is limited due to their unsupervised nature of the general sequence contexts they captured that is not specific to the protein being engineered. In this work, we propose ECNet, a deep-learning algorithm to exploit evolutionary contexts to predict functional fitness for protein engineering. Our method integrated local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest, as well as the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe. This biologically motivated sequence modeling approach enables accurate mapping from sequence to function and provides generalization from low-order mutants to higher-orders. Through extensive benchmark experiments, we showed that our method outperforms existing methods on [~]50 deep mutagenesis scanning and random mutagenesis datasets, demonstrating its potential of guiding and expediting protein engineering.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Inferring protein fitness landscapes from laboratory evolution experiments 97%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 96%
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 96%
Similar papers in this journal
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 96%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 95%
- Protein engineering via Bayesian optimization-guided evolutionary algorithm and robotic experiments 94%
Similar papers in this journal
- DynamicGT: a dynamic-aware geometric transformer model to predict protein binding interfaces in flexible and disordered regions 96%
- Sequence-based prediction of protein-protein interactions: a structure-aware interpretable deep learning model 95%
- Evolutionary velocity with protein language models 94%
Similar papers in this journal
- PIPENN: Protein Interface Prediction with an Ensemble of Neural Nets 96%
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 95%
- Deep Local Analysis deconstructs protein-protein interfaces and accurately estimates binding affinity changes upon mutation 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.