A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models
Sartori, J.; Guimaraes, A. C. R.; Machado, L. d. A.
Show abstract
Accurate computational prediction of enzyme function, standardized by Enzyme Commission (EC) numbers, is essential for large-scale genome annotation and generative enzyme design. However, it remains unclear whether state-of-the-art predictors learn the intrinsic structural determinants of catalytic activity or merely rely on global sequence similarity to annotated homologues. To address this gap, we introduce EnzymARC, a novel benchmark dataset of putative non-functional decoy sequences generated via structure-guided, systematic disruption of active sites (targeting catalytic residues and surrounding 5 A, 10 A, and 15 A radii) from experimentally annotated enzymes. We evaluated three distinct prediction paradigms against this dataset: homology-based annotation (DIAMOND), contrastive learning with protein language models (CLEAN), and a deep learning model incorporating non-enzyme discrimination (DeepEC). Our findings reveal that current models are highly vulnerable to phylogenetic shortcuts. Both DIAMOND and CLEAN exhibited false positive rates exceeding 90\% for low-perturbation decoys, confidently assigning the original EC numbers despite the destruction of the catalytic machinery. While DeepEC demonstrated improved sensitivity at higher perturbation levels, highlighting the benefit of negative training examples, all models struggled to identify targeted active-site disruptions. We demonstrate that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and we propose that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 94%
- Deducing high-accuracy protein contact-maps from a triplet of coevolutionary matrices through deep residual convolutional networks 93%
- Ig-VAE: Generative Modeling of Immunoglobulin Proteins by Direct 3D Coordinate Generation 92%
Similar papers in this journal
Similar papers in this journal
- ConforFold Recovers Alternative Protein Conformations Beyond MSA Subsampling 96%
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 94%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 94%
Similar papers in this journal
- Hybrid Deep Learning with Protein Language Models and Dual-Path Architecture for Predicting IDP Functions 95%
- Data-efficient protein mutational effect prediction with weak supervision by molecular simulation and protein language models 94%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.