Improve Protein Solubility and Activity based on Machine Learning Models
Han, X.; Ning, W.; Ma, X.; Wang, X.; Zhou, K.
Show abstract
Improving catalytic ability of protein biocatalysts leads to reduction in the production cost of biocatalytic manufacturing process, but the search space of possible proteins/mutants is too large to explore exhaustively through experiments. To some extent, highly soluble recombinant proteins tend to exhibit high activity. Here, we demonstrate that an optimization methodology based on machine learning prediction model can effectively predict which peptide tags can improve protein solubility quantitatively. Based on the protein sequence information, a support vector machine model we recently developed was used to evaluate protein solubility after randomly mutated tags were added to a target protein. The optimization algorithm guided the tags to evolve towards variants that can result in higher solubility. Moreover, the optimization results were validated successfully by adding the tags designed by our optimization algorithm to a model protein, expressing it in vivo and experimentally quantifying its solubility and activity. For example, solubility of a tyrosine ammonium lyase was more than doubled by adding two tags to its N- and C-terminus. Its protein activity was also increased nearly 3.5 fold by adding the tags. Additional experiments also supported that the designed tags were effective for improving activity of multiple proteins and are better than previously reported tags. The presented optimization methodology thus provides a valuable tool for understanding the correlation between amino acid sequence and protein solubility and for engineering protein biocatalysts.\n\nContactkang.zhou@nus.edu.sg, chewxia@nus.edu.sg
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Reengineering of a flavin-binding fluorescent protein using ProteinMPNN 93%
- Convergent behavior of extended stalk regions from staphylococcal surface proteins with widely divergent sequence patterns 92%
- Target-template relationships in protein structure prediction and their effect on the accuracy of thermostability calculations 92%
Similar papers in this journal
- Characterizing a New Fluorescent Protein for Low Limit of Detection Sensing in the Cell-Free System 93%
- Engineering transcription factor BmoR mutants for constructing multifunctional alcohol biosensors 93%
- Engineering the substrate specificity of toluene degrading enzyme XylM using biosensor XylS and machine learning 93%
Similar papers in this journal
- DeepSCM: an efficient convolutional neural network surrogate model for the screening of therapeutic antibody viscosity 92%
- DeepSP: Deep Learning-Based Spatial Properties to Predict Monoclonal Antibody Stability 92%
- Machine learning-assisted medium optimization revealed the discriminated strategies for improved production of the foreign and native metabolites 92%
Similar papers in this journal
- Single point mutations can potentially enhance infectivity of SARS-CoV-2 revealed by in silico affinity maturation and SPR assay 93%
- Mitoxantrone dihydrochloride, an FDA approved drug, binds with SARS-CoV-2 NSP1 C-terminal 91%
- Protein secondary structure prediction with context convolutional neural network 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.