Using experimental results of protein design to guide biomolecular energy-function development
Haddox, H. K.; Rocklin, G. J.; Motta, F. C.; Strickland, D.; Halabiya, S. F.; Cordray, C.; Park, H.; Klavins, E.; Baker, D.; DiMaio, F.
Show abstract
Computational models of macromolecules have many applications in biochemistry, but physical inaccuracies limit their utility. One class of models uses energy functions rooted in classical mechanics. The standard datasets used to train these models are limited in diversity, pointing to a need for new training data. Here, we sought to explore a new paradigm for training an energy function, where the Rosetta energy function was used to design de novo proteins. Experimental results on these designs were then used to identify failure modes of design, which were subsequently used as a "guiding principle" to retrain the energy function. Specifically, we examined a diverse set of de novo protein designs experimentally tested for their ability to stably fold, identifying unstable designs that were predicted to be stable by the Rosetta energy function. Using deep mutational scanning, we identified single amino-acid mutations that rescued the stability of these designs, providing insight into common failure modes of the energy function. We identified one key failure mode, involving steric clashing in protein cores. We identified similar overpacking when using Rosetta to refine high-resolution protein crystal structures, quantified the degree of overpacking, and refit a small set of energy-function parameters to better recapitulate native-like packing. Following fitting, we largely eliminated the failure mode in the refinement task, while retaining performance on other benchmarks, resulting in an updated version of the Rosetta energy function. This work shows how learning from protein designs can guide energy-function development. Author summaryComputational models of macromolecules have many applications, such as predicting structures, predicting mutational effects, or designing new proteins. One type of model uses an energy function to explicitly model the physical forces at play. These models are trained using experimental data. However, available training data are limited in number and diversity, prompting a need for new sources of training data. In this paper, we explore a new paradigm for training an energy function, which involves using the energy function to design de novo proteins and then learning from which designs succeed and which fail when experimentally tested in the lab. We used experimental data to learn common failure modes in design, identified examples of a failure mode in a high-quality benchmark, and then used this benchmark to retrain the energy function. Using this strategy, we identified and largely resolved a bias in the Rosetta energy function that involved energetically unfavorable steric clashes in protein cores. Overall, this work helps establish a framework for how learning from design can be used to guide the development of macromolecular models.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Suite of Designed Protein Cages Using Machine Learning Algorithms and Protein Fragment-Based Protocols 95%
- Protein language model embeddings for fast, accurate, alignment-free protein structure prediction 95%
- Modeling of protein conformational changes with Rosetta guided by limited experimental data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.