Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity
Zhang, S.; Xie, L.; Tiourine, D.; Xie, L.
Show abstract
Predicting protein-ligand binding processes, such as affinity and kinetics, is critical for accelerating drug discovery. However, many existing computational methods face key limitations, including insufficient integration of comprehensive databases, inadequate representation of protein structural dynamics, and incomplete modeling of microscale protein-ligand interactions. To address these challenges, we introduce ProMoNet, a sequence-based pre-training and fine-tuning framework to enhance protein-ligand binding process prediction. ProMoNet connects protein and molecular foundation models to expand data coverage and enhance diversity, and it integrates large-scale binding site pre-training with efficient fine-tuning for affinity and kinetics prediction. During pre-training, it effectively models microscale protein-ligand interactions and captures the dynamic nature of proteins, including binding site crypticity, without relying on 3-dimensional structural inputs. Notably, ProMoNets pre-training module surpasses or matches state-of-the-art structure-based methods in identifying exposed and cryptic binding sites. In the fine-tuning stage, it transfers pre-trained knowledge, achieving superior performance in affinity and kinetics prediction tasks with high computational efficiency. The combination of ProMoNets powerful modeling capabilities and demonstrated success across multiple tasks highlight its potential for broad applications in drug discovery.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 97%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 96%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 96%
Similar papers in this journal
- Fast protein structure searching using structure graph embeddings 96%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.