Quantified Dynamics-Property Relationships: Data-Efficient Protein Engineering with Machine Learning of Protein Dynamics
Burgin, T. E.
Show abstract
Machine learning has proven to be very powerful for predicting mutation effects in proteins, but the simplest approaches require a substantial amount of training data. Because experiments to collect training data are often expensive, time-consuming, and/or otherwise limited, alternatives that make good use of small amounts of data to guide protein engineering are of high potential value. One potential alternative to large-scale benchtop experiments for collecting training data is high-throughput molecular dynamics simulation; however, to date this source of data has been largely absent from the literature. Here, I introduce a new method for selecting desirable protein variants based on quantified relationships between a small number of experimentally determined labels and descriptors of their dynamic properties. These descriptors are provided by deep neural networks trained on data from molecular dynamics simulations of variants of the protein of interest. I demonstrate that this approach can obtain very highly optimized variants based on small amounts of experimental data, outperforming alternative supervised approaches to machine learning-guided directed evolution with the same amount of experimental data. Furthermore, I show that quantified dynamics-property relationships based on only a handful of experimentally labeled example sequences can be used to accurately predict the key residues that are most relevant to determining the property in question, even when that information could not have been known or predicted based on either the molecular dynamics simulations or the experimental data alone. This work establishes a new and practical framework for incorporating general protein dynamics information from simulations of mutants to guide protein engineering. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/650227v3_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@302943org.highwire.dtl.DTLVardef@1e4f709org.highwire.dtl.DTLVardef@1167c20org.highwire.dtl.DTLVardef@12f3627_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOToc GraphicC_FLOATNO C_FIG
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MELD-Adapt: On-the-Fly Belief Updating in Integrative Molecular Dynamics 97%
- SOP-MULTI: A self-organized polymer based coarse-grained model for multi-domain and intrinsically disordered proteins with conformation ensemble consistent with experimental scattering data 97%
- A linear response theory based method for prediction of large scale protein conformational changes upon ligand binding 97%
Similar papers in this journal
- Disentangling folding from energetic traps in simulations of disordered proteins 97%
- Machine learning of molecular dynamics simulations provides insights into modulation of viral capsid assembly 97%
- Searching for Structure: Characterizing the Protein Conformational Landscape with Clustering-based Algorithms 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- SMOG 2 and OpenSMOG: Extending the limits of structure-based models 96%
- Investigating the determinants of performance in machine learning for protein fitness prediction 95%
- Target-template relationships in protein structure prediction and their effect on the accuracy of thermostability calculations 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.