Machine learning models for the prediction of enzyme properties should be tested on proteins not used for model training
Kroll, A.; Lercher, M. J.
Show abstract
The recently published DLKcat model, a deep learning approach for predicting enzyme turnover numbers (kcat), claims to enable high-throughput kcat predictions for metabolic enzymes from any organism and to capture kcat changes for mutated enzymes. Here, we critically evaluate these claims. We show that DLKcat predictions become positively misleading for enzymes with less than 60% sequence identity to the training data, performing worse than simply assuming a mean kcat value for all reactions. Furthermore, DLKcats ability to predict mutation effects is much weaker than implied, capturing only 3% of the experimentally observed variation across mutants not included in the training data. These findings highlight significant limitations in DLKcats generalizability and its practical utility for predicting kcat values for novel enzyme families or mutants, which are crucial applications in fields such as metabolic modeling.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Designing diverse and high-performance proteins with a large language model in the loop 93%
- Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP) 93%
- Medusa: software to build and analyze ensembles of genome-scale metabolic network reconstructions 93%
Similar papers in this journal
Similar papers in this journal
- TemBERTure: Advancing protein thermostability prediction with Deep Learning and attention mechanisms 92%
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 92%
- Beyond synthetic lethality in large-scale metabolic and regulatory network models via genetic minimal intervention sets 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.