Prioritizing Stability-enhancing Mutations using a Protein Language Model in conjunction with Physics-Based Predictions
Rhodes, E. R.; Scarabelli, G.; Jou, J.; Byerly, J.; Sprenger, K. G.; Oloo, E. O.
Show abstract
Mutational protein engineering, as currently practiced in the biotechnology and pharmaceutical industries, is both tedious and expensive. Computationally driven protein design has the potential to expedite the process and generate high-quality variants at a lower cost. Using datasets of 174824 mutations for 180 proteins, we benchmark the effectiveness of a protein language model (PLM), Evolutionary Scale Modeling (ESM), alongside a physics-based method (PBM), namely molecular mechanics energies with generalized Born and surface area continuum solvation (MM/GBSA), as triaging tools for identifying and prioritizing target positions and specific mutations that improve protein thermodynamic stability. We found prediction biases in each method but also determined that these biases can be mitigated by applying the two methods in a complementary manner. We propose a hybrid mutation prioritization and selection strategy that achieves better accuracy than either method alone. Through re-ranking, the combined prioritization strategy attained a higher overall average ROC AUC of 0.744 across the dataset compared to either MM/GBSA alone (0.684) or ESM Log Odds alone (0.597).
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MCSS-based Predictions of Binding Mode and Selectivity of Nucleotide Ligands 95%
- AI-Based Methods for Cryptic Pocket Detection Are Fast and Qualitative Compared to Quantitatively Predictive Simulations 95%
- Thermal Adaptation of Cytosolic Malate Dehydrogenase Revealed by Deep Learning and Coevolutionary Analysis 95%
Similar papers in this journal
- Target-template relationships in protein structure prediction and their effect on the accuracy of thermostability calculations 96%
- Investigating the determinants of performance in machine learning for protein fitness prediction 96%
- Improved prediction of site-rates from structure with averaging across homologs 95%
Similar papers in this journal
- Disentangling the contribution of each descriptive characteristic of every single mutation to its functional effects 97%
- Disentangling folding from energetic traps in simulations of disordered proteins 95%
- ProAffinity-GNN: A Novel Approach to Structure-based Protein-Protein Binding Affinity Prediction via a Curated Dataset and Graph Neural Networks 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.