A Two-Stage ESM-Based Machine Learning Pipeline for Robust Hierarchical Enzyme Function Prediction
Hua, X.; Grimaud, G. M.
Show abstract
Accurate enzyme annotation remains a major bottleneck in translating rapidly growing protein sequence data into biological knowledge. Enzyme Commission (EC) prediction is particularly challenging because enzyme functions are organized hierarchically, annotations are often imbalanced across classes, and sequence similarity alone may be insufficient to resolve functional differences. To address these challenges, we developed ESM-ECForest, a two-stage framework that combines protein embeddings generated by the pretrained language model ESM-2 (Evolutionary Scale Modeling 2) with Random Forest classifiers. The first stage distinguishes enzymes from non-enzymes, whereas the second assigns one or more EC numbers to proteins predicted to be enzymatic. On an external benchmark comprising 25,778 protein sequences, ESM-ECForest achieved the highest weighted F1 score among the evaluated methods at all four EC levels, decreasing from 0.94 at Level 1 to 0.90 at Level 4. The largest relative improvements were observed for lyases (EC 4), ligases (EC 6), and translocases (EC 7), although EC 6 and EC 7 remained the most difficult classes internally. Visualization of the ESM-2 embedding space using Uniform Manifold Approximation and Projection (UMAP) revealed clustering patterns consistent with enzyme functional relationships, indicating that biologically relevant information is retained in the pretrained representations prior to supervised classification. These results support the use of pretrained protein language model embeddings as an effective foundation for enzyme annotation. By combining large-scale sequence representations with a lightweight supervised classifier, ESM-ECForest provides a scalable approach for EC prediction and may facilitate functional annotation of protein sequences derived from large genomic and metagenomic datasets.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TALE: Transformer-based protein function Annotation with joint sequence-Label Embedding 94%
- FAPM: Functional Annotation of Proteins using Multi-Modal Models Beyond Structural Modeling 94%
- Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function 94%
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 93%
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 92%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 92%
Similar papers in this journal
- Metabolic pathway prediction using non-negative matrix factorization with improved precision 92%
- Critiquing Protein Family Classification Models Using Sufficient Input Subsets 92%
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.