DeepRES: Deep learning enables reaction-based comprehensive enzyme screening
Hirota, K.; Yamada, T.
Show abstract
BackgroundEnzymes accelerate biochemical reactions in living organisms, thus playing an important role in metabolism. Although metabolic pathway databases are growing, many metabolic reactions, termed orphan enzymes, have not been annotated to gene sequences, which hinders functional annotation in genomic analysis. Moreover, protein databases contain many proteins of unknown function. Owing to this gap between known proteins and enzymatic reactions, various proteins of unknown function may be orphan enzymes; however, available tools cannot adequately predict these links. ResultsIn this study, we developed DeepRES, an AI-based framework for comprehensive enzyme screening, to explore novel enzyme candidates from proteins of unknown function for reactions of interest. DeepRES implements enzyme screening via two steps: classification of enzymes and non-enzymes and prediction of catalytic capabilities for enzyme-reaction pairs. The two deep learning models comprising DeepRES showed comparable or superior performance to that of existing software. We performed screening of 1,255 orphan enzymes involved in the microbiome using DeepRES and successfully identified candidate proteins for 897 orphan enzymes. We then used those candidates as references for genomic analysis and explored novel biosynthetic gene clusters from microbial genomes to obtain promising candidate gene clusters, including those related to anthocyanin degradation. ConclusionsComprehensive enzyme screening via DeepRES, which is the first computational tool designed to associate orphan enzymes with proteins of unknown function, is expected to facilitate high-throughput identification of orphan enzyme-encoding genes. Furthermore, DeepRES can be easily integrated into the current genomic analysis pipeline to extend the functional annotation.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An Analysis of Protein Language Model Embeddings for Fold Prediction 96%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 96%
- Seq2Topt: a sequence-based deep learning predictor of enzyme optimal temperature 96%
Similar papers in this journal
- iPRESTO: automated discovery of biosynthetic sub-clusters linked to specific natural product substructures 96%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 95%
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 94%
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 96%
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 95%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.