Using Autoregressive-Transformer Model for Protein-Ligand Binding Site Prediction
Pourmirzaei, M.; Alqarghuli, S.; Esmaili, F.; POURMIRZAEIOLIAEI, M.; Rezaei, M.; Xu, D.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWAccurate prediction of protein-ligand binding sites is critical for understanding molecular interactions and advancing drug discovery. Existing computational approaches often suffer from limited generality, restricting their applicability to a small subset of ligands, while data scarcity further impairs performance, particularly for underrepresented ligand types. To address these challenges, we introduce a unified model that integrates a protein language model with an autoregressive transformer for protein-ligand binding site prediction. By framing the task as a language modeling problem and incorporating task-specific tokens, our method achieves broad ligand coverage while relying solely on protein sequence input. We systematically analyze ligand-specific task token embeddings, demonstrating that they capture meaningful biochemical properties through clustering and correlation analyses. Furthermore, our multi-task learning strategy enables effective knowledge transfer across ligands, significantly improving predictions for those with limited training data. Experimental evaluations on 41 ligands highlight the models superior generalization and applicability compared to existing methods. This work establishes a scalable generative AI framework for binding site prediction, laying the foundation for future extensions incorporating structural information and richer ligand representations. The code, model, and datasets are available at this link.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- All-Atom Protein Sequence Design using Discrete Diffusion Models 96%
- Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity 96%
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 95%
Similar papers in this journal
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 96%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 96%
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 96%
Similar papers in this journal
- Dual-channel graph learning reveals similarity and complementarity in protein-protein interaction networks 95%
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 95%
- Designing diverse and high-performance proteins with a large language model in the loop 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.