Back

KcatNet: advancing genome-wide enzyme turnover number prediction through structural enzymatic characterization

Pan, T.; Cui, X.; Koh, H. Y.; Bi, Y.; Wang, X.; Zhang, Y.; Hu, S.; Webb, G. I.; Gasser, R.; Kurgan, L.; Zhang, G.; Song, J.

2025-03-13 bioinformatics
10.1101/2025.03.09.642294 bioRxiv
Show abstract

Enzyme turnover numbers (Kcat) are fundamental kinetic constants that quantify enzymatic efficiency. Systematic studies of Kcat are essential for characterizing the mechanisms underlying proteomic composition and cellular metabolism. However, experimental measurements of Kcat remain limited and prone to noise. To address this, we present KcatNet, a geometric deep learning model designed for high-throughput prediction of Kcat in metabolic enzymes across all organisms, leveraging paired enzyme sequence and substrate representations. KcatNet consistently outperforms existing predictors, particularly for enzymes with high catalytic efficiency, and demonstrates strong generalization to enzymes that are dissimilar to those in the training set. Furthermore, KcatNet uncovers structural mechanisms and interaction patterns within enzyme-substrate complexes, providing insights into architectural principles that are inaccessible with existing methods by harnessing the representational power of large-scale protein language models. We applied KcatNet to genome-scale Kcat prediction across diverse yeast species, improving proteome allocation predictions by integrating its outputs into metabolic models. Experimental validation further confirmed the models ability to identify enzyme mutants with enhanced activity. By bridging the gap between sequence, structure, and function, KcatNet establishes a robust foundation for advancing understanding of molecular-level mechanisms and accelerating enzyme engineering efforts.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.