Integration of Machine Learning Improves the Prediction Accuracy of Molecular Modelling for M. jannaschii Tyrosyl-tRNA Synthetase Substrate Specificity
Duan, B.; Sun, Y.
Show abstract
Design of enzyme binding pocket to accommodate substrates with different chemical structure is a great challenge. Traditionally, thousands even millions of mutants have to be screened in wet-lab experiment to find a ligand-specific mutant and large amount of time and resources is consumed. To accelerate the screening process, here we propose a novel workflow through integration of molecular modeling and data-driven machine learning method to generate mutant libraries with high enrichment ratio for recognition of specific substrate. M. jannaschii tyrosyl-tRNA synthetase (Mj. TyrRS) is used as an example system to give a proof of concept since the sequence and structure of many unnatural amino acid specific Mj. TyrRS mutants have been reported. Based on the crystal structures of different Mj. TyrRS mutants and Rosetta modeling result, we find D158G/P is the critical residue which influences the backbone disruption of helix with residue 158-163. Our results show that compared with random mutation, Rosetta modeling and score function calculation can elevate the enrichment ratio of desired mutants by 2-fold in a test library having 687 mutants, while after calibration by machine learning model trained using known data of Mj. TyrRS mutants and ligand, the enrichment ratio can be elevated by 11-fold. This molecular modeling and machine learning-integrated workflow is anticipated to significantly benefit to the Mj. tyrRS mutant screening and substantially reduce the time and cost of web-lab experiment. Besides, this novel process will have broad application in the field of computational protein design. CCS Concepts* Applied computing * Life and medical sciences * Computational biology * Molecular structural biology
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- De novo drug designing coupled with brute force screening and structure guided lead optimization gives highly specific inhibitor of METTL3: a potential cure for Acute Myeloid Leukaemia 97%
- A strategy to optimize the peptide-based inhibitors against different mutants of the spike protein of SARS-CoV-2 97%
- Dynamic conformational states of apo and cabozantinib bound TAM kinases to differentiate active-inactive kinetic models 96%
Similar papers in this journal
- Identification of Family-Specific Features in Cas9 and Cas12 Proteins: A Machine Learning Approach Using Complete Protein Feature Spectrum 97%
- Structural characterization of LsrK to target quorum sensing and comparison between X-ray and homology model 96%
- Streamlining Computational Fragment-Based Drug Discovery through Evolutionary Optimization Informed by Ligand-Based Virtual Prescreening 96%
Similar papers in this journal
- Molecular Glue-Design-Evaluator (MOLDE): An Advanced Method for In-Silico Molecular Glue Design 97%
- Utilizing Heteroatom Types and Numbers from Extensive Ligand Libraries to Develop Novel hERG Blocker QSAR Models Using Machine Learning-based Classifiers 96%
- Benchmarking HelixFold3-Predicted Holo Structures for Relative Free Energy Perturbation Calculations 95%
Similar papers in this journal
- Mitoxantrone dihydrochloride, an FDA approved drug, binds with SARS-CoV-2 NSP1 C-terminal 96%
- Single point mutations can potentially enhance infectivity of SARS-CoV-2 revealed by in silico affinity maturation and SPR assay 96%
- Protein secondary structure prediction with context convolutional neural network 92%
Similar papers in this journal
- Mechanistic insights into the Japanese Encephalitis Virus RNA dependent RNA polymerase protein inhibition by bioflavonoids from Azadirachta indica 97%
- In Silico Analysis Predicting Effects of Deleterious SNPs of Human RASSF5 Gene on its Structure and Functions 96%
- Mechanistic insights into the deleterious role of nasu-hakola disease associated TREM2 variants 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.