Back

A Genetic Algorithm to scour protein sequence space: A novel framework for protein engineering using protein language models and force fields

Sartori, J.; Krempser, E.; Guimaraes, A. C. R.; Machado, L. d. A.

2025-10-31 bioinformatics
10.1101/2025.10.29.685431 bioRxiv
Show abstract

The optimization of protein sequences for enhanced binding and stability remains a formidable challenge in bioengineering due to the vastness of sequence space. Existing state-of-the-art methods, including traditional structure-based design and protein language models, use fitness estimators as objective functions to guide search algorithms that scour sequence space for optimal sequences. However, effective exploration requires search strategies that balance diversity and computational efficiency, as both are paramount for effective exploration. This work presents GAPO (Genetic Algorithm for Protein Optimization), a novel flexible framework that integrates evolutionary computing, protein language models, and structure-based design to efficiently explore sequence space. GAPO employs genetic algorithms to iteratively refine protein sequences based on user-defined objective functions, using either force field-derived information or protein language models to assess fitness, it also allows users to define custom objective functions, including multi-objective ones. We detail GAPOs implementation, highlighting features such as customizable initialization methods, diverse selection strategies, and mutation techniques informed by evolutionary scale modeling (ESM).IIn a case study using Hen-egg lysozyme, GAPO outperformed simulated annealing (SA) in both protein language model and energy-based objectives, converging to higher average ESM2 probabilities (0.98 vs. WT 0.89 and SA 0.88) and more favorable REF2015 energies (-510 REU vs. WT -415 REU and SA -405 REU), while maintaining reproducible behavior across independent runs. GAPO is available at https://github.com/izzetbiophysicist/GAPO.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.