The substrate scopes of enzymes: a general prediction model based on machine and deep learning
Kroll, A.; Ranjan, S.; Engqvist, M. K.; Lercher, M. J.
Show abstract
For a comprehensive understanding of metabolism, it is necessary to know all potential substrates for each enzyme encoded in an organisms genome. However, for most proteins annotated as enzymes, it is unknown which primary and/or secondary reactions they catalyze [1], as experimental characterizations are time-consuming and costly. Machine learning predictions could provide an efficient alternative, but are hampered by a lack of information regarding enzyme non-substrates, as available training data comprises mainly positive examples. Here, we present ESP, a general machine learning model for the prediction of enzyme-substrate pairs, with an accuracy of over 90% on independent and diverse test data. This accuracy was achieved by representing enzymes through a modified transformer model [2] with a trained, task-specific token, and by augmenting the positive training data by randomly sampling small molecules and assigning them as non-substrates. ESP can be applied successfully across widely different enzymes and a broad range of metabolites. It outperforms recently published models designed for individual, well-studied enzyme families, which use much more detailed input data [3, 4]. We implemented a user-friendly web server to predict the substrate scope of arbitrary enzymes, which may support not only basic science, but also the development of pharmaceuticals and bioengineering processes.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP) 95%
- Knowledge-guided data mining on the standardized architecture of NRPS: subtypes, novel motifs, and sequence entanglements 94%
- Transfer learning enables prediction of CYP2D6 haplotype function 94%
Similar papers in this journal
- Adding stochastic negative examples into machine learning improves molecular bioactivity prediction 95%
- CENsible: Interpretable Insights into Small-Molecule Binding with Context Explanation Networks 95%
- DrugHIVE: Target-specific spatial drug design and optimization with a hierarchical generative model 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.