Pretrained language models and weight redistributionachieve precise kcat prediction
Yu, H.; Luo, X.
Show abstract
The enzyme turnover number (kcat) is a meaningful and valuable kinetic parameter, reflecting the catalytic efficiency of an enzyme to a specific substrate, which determines the global proteome allocation, metabolic fluxes and cell growth. Here, we present a precise kcat prediction model (PreKcat) leveraging pretrained language models and a weight redistribution strategy. PreKcat significantly outperforms the previous kcat prediction method in terms of various evaluation metrics. We also confirmed the ability of PreKcat to discriminate enzymes of different metabolic contexts and different types. Additionally, the proposed weight redistribution strategies effectively reduce the prediction error of high kcat values and capture minor effects of amino acid substitutions on two crucial enzymes of the naringenin synthetic pathway, leading to obvious distinctions. Overall, the presented kcat prediction model provides a valuable tool for deciphering the mechanisms of enzyme kinetics and enables novel insights into enzymology and biomedical applications.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning 94%
- Knowledge-guided data mining on the standardized architecture of NRPS: subtypes, novel motifs, and sequence entanglements 94%
- Explainable deep transfer learning model for disease risk prediction using high-dimensional genomic data 94%
Similar papers in this journal
- Seq2Topt: a sequence-based deep learning predictor of enzyme optimal temperature 97%
- Enhanced compound-protein binding affinity prediction by representing protein multimodal information via a coevolutionary strategy 97%
- AI-Guided Discovery and Optimization of Antimicrobial Peptides Through Species-Aware Language Model 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.