Heuristic Multi-site Optimization for Protein Sequence Design using Masked Protein Language Models
Wang, L.; Wang, Y.; Qiu, C.; Xiao, L.; Liu, X.; Chen, J.
Show abstract
Protein sequence design for tailored functional properties is a fundamental task in protein engineering, with critical applications in drug discovery and therapeutic development. Efficient navigation of the combinatorial vastness of protein sequence space to identify functional variants remains a formidable challenge. Conventional approaches, which predominantly rely on template-based local search or single-residue mutagenesis, are constrained by their susceptibility to local optima and their potential risk of destabilizing native structural stability. In this study, we introduce ProtHMSO, a heuristic multi-site optimization framework leveraging masked protein language models (ProtLMs) for context-aware sequence exploration. ProtHMSO mimics natural evolutionary mechanisms by employing ProtLM-derived substitution probabilities to guide heuristic searches for synergistic mutations, thereby constraining combinatorial search spaces through evolutionary and biophysical priors. Given a starting sequence and a predictive fitness model, ProtHMSO heuristically generates high-fitness multi-site variants under evolutionary constraints. Besides, ProtHMSO can also be as a modular plugin that enhances the convergence efficiency of genetic algorithms (GAs) and Monte Carlo tree search (MCTS) by integrating evolutionary priors into exploration strategies. Benchmark experiments demonstrate that protein sequences generated by ProtHMSO exhibit superior functional performance and closer alignment with natural sequence distribution, compared with state-of-the-art methods. These advancements highlight that ProtHMSO has strong potential and compatibility to accelerate functional protein discovery, offering a robust framework for efficient and context-aware exploration of protein sequence space. Author summaryTo address the challenge of efficiently discovering functional new proteins in protein engineering due to the vast sequence space, and to overcome the limitations of traditional evolutionary algorithms that rely on blind random mutagenesis, resulting in inefficiency and prone to structural destabilization, we proposed a heuristic multi-site optimization framework, ProtHMSO. Its core concept is to leverage the powerful contextual prediction capabilities of masked protein language models (such as ESM-2) to guide sequence mutagenesis. By predicting amino acid substitutions at specific sites that are consistent with evolutionary laws and biophysical priors, ProtHMSO narrows the exploration scope from the vast combinatorial space to a small number of high-potential candidate sequences, achieving intelligent and efficient optimization of protein sequences. Furthermore, ProtHMSO is not just a standalone algorithm, but also a plug-and-play "enhancement module". By embedding it into a genetic algorithm (GA) and a Monte Carlo tree search (MCTS), it replaces the random mutation operator in the former with its intelligent mutation and guides the tree expansion process in the latter. This enables these classic optimization algorithms to break free from the blindness of exploration and achieve faster convergence and better results, demonstrating the wide applicability and great potential of this framework in improving the performance of tools in the entire field of computational protein design.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Novel, provable algorithms for efficient ensemble-based computational protein design and their application to the redesign of the c-Raf-RBD:KRas protein-protein interface 95%
- Designing diverse and high-performance proteins with a large language model in the loop 94%
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 94%
Similar papers in this journal
Similar papers in this journal
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 96%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network 96%
Similar papers in this journal
- Physical-aware model accuracy estimation for protein complex using deep learning method 95%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 95%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.