Back

Machine learning models for the prediction of enzyme properties should be tested on proteins not used for model training

Kroll, A.; Lercher, M. J.

2023-02-07 bioinformatics
10.1101/2023.02.06.526991 bioRxiv
Show abstract

The recently published DLKcat model, a deep learning approach for predicting enzyme turnover numbers (kcat), claims to enable high-throughput kcat predictions for metabolic enzymes from any organism and to capture kcat changes for mutated enzymes. Here, we critically evaluate these claims. We show that DLKcat predictions become positively misleading for enzymes with less than 60% sequence identity to the training data, performing worse than simply assuming a mean kcat value for all reactions. Furthermore, DLKcats ability to predict mutation effects is much weaker than implied, capturing only 3% of the experimentally observed variation across mutants not included in the training data. These findings highlight significant limitations in DLKcats generalizability and its practical utility for predicting kcat values for novel enzyme families or mutants, which are crucial applications in fields such as metabolic modeling.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.