CPI-Pred: A deep learning framework for predicting functional parameters of compound-protein interactions
Xu, Z.; Barghout, R. A.; Wu, J.; Garg, D.; Song, Y. S.; Mahadevan, R.
Show abstract
Recent advancements in deep learning have enabled functional annotation of genome sequences, facilitating the discovery of new enzymes and metabolites. However, accurately predicting compound-protein interactions (CPI) from sequences remains challenging due to the complexity of these interactions and the sparsity and heterogeneity of available data, which constrain the generalization of patterns across their solution space. In this work, we introduce CPI-Pred, a versatile deep learning model designed to predict compound-protein interaction function. CPI-Pred integrates compound representations derived from a novel message-passing neural network and enzyme representations generated by state-of-the-art protein language models, leveraging innovative sequence pooling and cross-attention mechanisms. To train and evaluate CPI-Pred, we compiled the largest dataset of enzyme kinetic parameters to date, encompassing four key metrics: the Michaelis-Menten constant (KM), enzyme turnover number (kcat), catalytic efficiency (kcat/KM), and inhibition constant (KI).These kinetic parameters are critical for elucidating enzyme function in metabolic contexts and understanding their regulation by compounds within biological networks. We demonstrate that CPI-Pred can predict diverse types of CPI using only the amino acid sequence of enzymes and structural representations of compounds, outperforming state-of-the-art models on unseen compounds and structurally dissimilar enzymes. Over workflow provides a valuable tool for tackling a range of metabolic engineering challenges, including the designing of novel enzyme sequences and compounds, such as enzyme inhibitors. Additionally, the datasets curated in this study offer a valuable resource for the scientific community, serving as a benchmark for machine learning models focused on enzyme activity and promiscuity prediction.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ProAffinity-GNN: A Novel Approach to Structure-based Protein-Protein Binding Affinity Prediction via a Curated Dataset and Graph Neural Networks 96%
- DENVIS: scalable and high-throughput virtual screening using graph neural networks with atomic and surface protein pocket features 96%
- Graph Attention Site Prediction (GrASP): Identifying Druggable Binding Sites Using Graph Neural Networks with Attention 96%
Similar papers in this journal
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 96%
- Deep learning molecular interaction motifs from Receptor structure alone 95%
- Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data 95%
Similar papers in this journal
Similar papers in this journal
- DrugTar Improves Druggability Prediction by Integrating Large Language Models and Gene Ontologies 96%
- Attention-based approach to predict drug-target interactions across seven target superfamilies 96%
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 96%
Similar papers in this journal
- Deep Learning for Protein Peptide bindingPrediction: Incorporating Sequence, Structural andLanguage Model Features 95%
- Thinking like a structural biologist: A pocket-based 3D molecule generative model fueled by electron density 95%
- PIPENN-EMB: ensemble net and protein embeddings generalise protein interface prediction beyond homology 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.