Predicting enzyme substrate chemical structure with protein language models
Jinich, A.; Nazia, S. Z.; Tellez, A. V.; Rappoport, D.; Rhee, K. Y.
Show abstract
The number of unannotated or orphan enzymes vastly outnumber those for which the chemical structure of the substrates are known. While a number of enzyme function prediction algorithms exist, these often predict Enzyme Commission (EC) numbers or enzyme family, which limits their ability to generate experimentally testable hypotheses. Here, we harness protein language models, cheminformatics, and machine learning classification techniques to accelerate the annotation of orphan enzymes by predicting their substrates chemical structural class. We use the orphan enzymes of Mycobacterium tuberculosis as a case study, focusing on two protein families that are highly abundant in its proteome: the short-chain dehydrogenase/reductases (SDRs) and the S-adenosylmethionine (SAM)-dependent methyltransferases. Training machine learning classification models that take as input the protein sequence embeddings obtained from a pre-trained, self-supervised protein language model results in excellent accuracy for a wide variety of prediction tasks. These include redox cofactor preference for SDRs; small-molecule vs. polymer (i.e. protein, DNA or RNA) substrate preference for SAM-dependent methyltransferases; as well as more detailed chemical structural predictions for the preferred substrates of both enzyme families. We then use these trained classifiers to generate predictions for the full set of unannotated SDRs and SAM-methyltransferases in the proteomes of M. tuberculosis and other mycobacteria, generating a set of biochemically testable hypotheses. Our approach can be extended and generalized to other enzyme families and organisms, and we envision it will help accelerate the annotation of a large number of orphan enzymes. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=106 SRC="FIGDIR/small/509940v3_ufig1.gif" ALT="Figure 1"> View larger version (20K): org.highwire.dtl.DTLVardef@1dab89forg.highwire.dtl.DTLVardef@8eec65org.highwire.dtl.DTLVardef@141ff13org.highwire.dtl.DTLVardef@1d16212_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 95%
- Generalizable and scalable protein stability prediction with rewired protein generative models 95%
- Accuracy and data efficiency in deep learning models of protein expression 95%
Similar papers in this journal
- SHARK enables homology assessment in unalignable anddisordered sequences 96%
- Integrated evolutionary and structural analysis reveals xenobiotics and pathogens as the major drivers of mammalian adaptation 95%
- Parametrically guided design of beta barrels and transmembrane nanopores using deep learning 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.