Back

Machine learning assisted ligand binding energy prediction for in silico generated glycosyl hydrolase enzyme combinatorial mutant library

Guranovic, I.; Kumar, M.; Bandi, C. K.; Chundawat, S. P. S.

2022-12-02 bioengineering
10.1101/2022.11.29.518414 bioRxiv
Show abstract

Molecular docking is a computational method used to predict the preferred binding orientation of one molecule to another when bound to each other to form an energetically stable complex. This approach has been widely used for early-stage small-molecule drug design as well as identifying suitable protein-based macromolecule residues for mutagenesis. Estimating binding free energy, based on docking interactions of protein to its ligand based on an appropriate scoring function is often critical for protein mutagenesis studies to improve the activity or alter the specificity of targeted enzymes. However, calculating docking free energy for a large number of protein mutants is computationally challenging and time-consuming. Here, we showcase an end-to-end computational workflow for predicting the binding energy of pNP-Xylose substrate docked within the substrate binding site for a large library of combinatorial mutants of an alpha-L-fucosidase (TmAfc, PDB ID-2ZWY) belonging to Thermotoga maritima glycosyl hydrolase (GH) family 29. Briefly, in silico combinatorial mutagenesis was performed for the top conserved residues in TmAfc as determined by running multiple sequence alignment against all GH29 family enzyme sequences downloaded from an in-house developed Carbohydrate-Active enZyme (CAZy) database retriever program. The binding energy was calculated through Autodock Vina with pNP-Xylose ligand docking with energy minimized TmAfc mutants, and the data was then used to train a neural network model which was also validated for model predictions using data from Autodock Vina. The current workflow can be adopted for any family of CAZymes to rapidly identify the effect of different mutations within the active site on substrate binding free energy to identify suitable targets for mutagenesis. We anticipate that this workflow could also serve as the starting point for performing more sophisticated and computationally intensive binding free energy calculations to identify targets for mutagenesis and hence optimize use of wet lab resources.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
Glycobiology
35 papers in training set
Top 0.1%
8.8%
2
Applied Microbiology and Biotechnology
32 papers in training set
Top 0.1%
6.7%
3
Biotechnology and Bioengineering
53 papers in training set
Top 0.1%
6.2%
4
Frontiers in Chemistry
16 papers in training set
Top 0.1%
6.2%
5
International Journal of Biological Macromolecules
76 papers in training set
Top 0.2%
6.2%
6
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 0.2%
5.1%
7
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.8%
4.4%
8
PLOS ONE
5266 papers in training set
Top 35%
4.0%
9
ACS Omega
105 papers in training set
Top 0.6%
3.2%
50% of probability mass above
10
Scientific Reports
3612 papers in training set
Top 35%
3.2%
11
Protein Engineering, Design and Selection
15 papers in training set
Top 0.1%
3.2%
12
Journal of Biomolecular Structure and Dynamics
43 papers in training set
Top 0.5%
2.8%
13
PeerJ
308 papers in training set
Top 4%
2.4%
14
Frontiers in Molecular Biosciences
102 papers in training set
Top 0.4%
2.4%
15
Journal of Cheminformatics
29 papers in training set
Top 0.3%
2.1%
16
Protein Science
246 papers in training set
Top 2%
2.1%
17
International Journal of Molecular Sciences
494 papers in training set
Top 6%
2.1%
18
PLOS Computational Biology
1863 papers in training set
Top 13%
2.1%
19
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.7%
1.9%
20
Journal of Biosciences
15 papers in training set
Top 0.1%
1.7%
21
eLife
5828 papers in training set
Top 51%
1.7%
22
Journal of Biotechnology
11 papers in training set
Top 0.1%
1.4%
23
Computers in Biology and Medicine
128 papers in training set
Top 3%
1.1%
24
Biochemical Journal
91 papers in training set
Top 1%
1.1%
25
Frontiers in Pharmacology
111 papers in training set
Top 3%
1.0%
26
Pharmaceuticals
34 papers in training set
Top 1%
0.8%
27
Bioscience Reports
27 papers in training set
Top 1%
0.8%
28
Biochimie
25 papers in training set
Top 0.8%
0.8%
29
Plant Physiology and Biochemistry
20 papers in training set
Top 1%
0.6%
30
Biomolecules
100 papers in training set
Top 4%
0.6%