CatPred: A comprehensive framework for deep learning in vitro enzyme kinetic parameters kcat, Km and Ki
Boorla, V. S.; Maranas, C. D.
Show abstract
Quantification of enzymatic activities still heavily relies on experimental assays, which can be expensive and time-consuming. Therefore, methods that enable accurate predictions of enzyme activity can serve as effective digital twins. A few recent studies have shown the possibility of training machine learning (ML) models for predicting the enzyme turnover numbers (kcat) and Michaelis constants (Km) using only features derived from enzyme sequences and substrate chemical topologies by training on in vitro measurements. However, several challenges remain such as lack of standardized training datasets, evaluation of predictive performance on out-of-distribution examples, and model uncertainty quantification. Here, we introduce CatPred, a comprehensive framework for ML prediction of in vitro enzyme kinetics. We explored different learning architectures and feature representations for enzymes including those utilizing pretrained protein language model features and pretrained three-dimensional structural features. We systematically evaluate the performance of trained models for predicting kcat, Km, and inhibition constants (Ki) of enzymatic reactions on held-out test sets with a special emphasis on out-of-distribution test samples (corresponding to enzyme sequences dissimilar from those encountered during training). CatPred assumes a probabilistic regression approach offering query-specific standard deviation and mean value predictions. Results on unseen data confirm that accuracy in enzyme parameter predictions made by CatPred positively correlate with lower predicted variances. Incorporating pre-trained language model features is found to be enabling for achieving robust performance on out-of-distribution samples. Test evaluations on both held-out and out-of-distribution test datasets confirm that CatPred performs at least competitively with existing methods while simultaneously offering robust uncertainty quantification. CatPred offers wider scope and larger data coverage ([~]23k, 41k, 12k data-points respectively for kcat, Km and Ki). A web-resource to use the trained models is made available at: https://tiny.cc/catpred
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BioStructNet: Structure-Based Network with Transfer Learning for Predicting Biocatalyst Functions 96%
- KaMLs for Predicting Protein pKa Values and Ionization States: Are Trees All You Need? 96%
- Accurate Predictions of Molecular Properties of Proteins via Graph Neural Networks and Transfer Learning 96%
Similar papers in this journal
- PROTACable is an Integrative Computational Pipeline of 3-D Modeling and Deep Learning to Automate the De Novo Design of PROTACs 95%
- CENsible: Interpretable Insights into Small-Molecule Binding with Context Explanation Networks 95%
- Protein Engineering with Lightweight Graph Denoising Neural Networks 94%
Similar papers in this journal
- Sitetack: A Deep Learning Model that Improves PTM Predictionby Using Known PTMs 96%
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 96%
- Guiding Discovery of Protein Sequence-Structure-Function Modeling 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.