Predicting turnover number fold-changes to recover true mutation effects and overcome biases in mutant datasets
Rousset, Y.; Kroll, A.; Lercher, M.
Show abstract
Accurately predicting enzyme turnover rates (kcat) for variant enzymes is essential for understanding the functional consequences of genetic variation and for rational protein engineering, as experimental determination remains challenging. Current computational methods offer general models that predict kcat across diverse enzymes, but these approaches are limited by two key biases: (i) kcat values for a wild-type and its variants usually lie within one to two orders of magnitude, making naive wild-type-like predictions for all variants appear deceptively accurate when evaluations include different enzyme families, and (ii) variant measurements for the same enzyme often share a directional bias, i.e., they tend to consistently either increase or decrease catalytic activity relative to the wild-type kcat. Here, we present FCKcat, the first mutation-sensitive machine-learning framework that predicts fold changes in kcat between any two enzyme variants. Leveraging a large experimental dataset of enzyme variant measurements, FCKcat encodes both structural and functional sequence differences to capture true mutational effects without memorizing average values of training variants. On unseen mutants, it achieves an R2 of 0.51 for fold-change prediction and 0.72 for absolute kcat. On a common test dataset comprising unseen mutants, FCKcat outperforms all existing methods, correctly predicting the direction of kcat change in over 80% of cases. This work establishes a foundation for genuinely mutation-aware enzyme modeling and represents a crucial step toward predictive computational tools for enzyme design, biotechnology, and synthetic biology.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accuracy and data efficiency in deep learning models of protein expression 96%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 96%
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.