DrugGPT: A GPT-based Strategy for Designing Potential Ligands Targeting Specific Proteins
Li, Y.; Gao, C.; Song, X.; Wang, X.; Xu, Y.; Han, S.
Show abstract
DrugGPT presents a ligand design strategy based on the autoregressive model, GPT, focusing on chemical space exploration and the discovery of ligands for specific proteins. Deep learning language models have shown significant potential in various domains including protein design and biomedical text analysis, providing strong support for the proposition of DrugGPT. In this study, we employ the DrugGPT model to learn a substantial amount of protein-ligand binding data, aiming to discover novel molecules that can bind with specific proteins. This strategy not only significantly improves the efficiency of ligand design but also offers a swift and effective avenue for the drug development process, bringing new possibilities to the pharmaceutical domain. In our research, we particularly optimized and trained the GPT-2 model to better adapt to the requirements of drug design. Given the characteristics of proteins and ligands, we redesigned the tokenizer using the BPE algorithm, abandoned the original tokenizer, and trained the GPT-2 model from scratch. This improvement enables DrugGPT to more accurately capture and understand the structural information and chemical rules of drug molecules. It also enhances its comprehension of binding information between proteins and ligands, thereby generating potentially active drug candidate molecules. Theoretically, DrugGPT has significant advantages. During the model training process, DrugGPT aims to maximize the conditional probability and employs the back-propagation algorithm for training, making the training process more stable and avoiding the Mode Collapse problem that may occur in Generative Adversarial Networks in drug design. Furthermore, the design philosophy of DrugGPT endows it with strong generalization capabilities, giving it the potential to adapt to different tasks. In conclusion, DrugGPT provides a forward-thinking and practical new approach to ligand design. By optimizing the tokenizer and retraining the GPT-2 model, the ligand design process becomes more direct and efficient. This not only reflects the theoretical advantages of DrugGPT but also reveals its potential applications in the drug development process, thereby opening new perspectives and possibilities in the pharmaceutical field.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Retro Drug Design: From Target Properties to Molecular Structures 98%
- Streamlining Computational Fragment-Based Drug Discovery through Evolutionary Optimization Informed by Ligand-Based Virtual Prescreening 96%
- AutoLead: An LLM-Guided Bayesian Optimization Framework for Multi-Objective Lead Optimization 96%
Similar papers in this journal
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 97%
- Deep learning integration of molecular and interactome data for protein-compound interaction prediction 96%
- PL-PatchSurfer3: Improved Structure-Based Virtual Screening for Structure Variation Using 3D Zernike Descriptors 96%
Similar papers in this journal
- DrugForm-DTA: Towards real-world drug-target binding Affinity Model 96%
- Network-based estimation of therapeutic efficacy and adverse reaction potential for prioritisation of anti-cancer drug combinations 94%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 94%
Similar papers in this journal
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 96%
- A Graph-Attention-Based Deep Learning Network for Predicting Biotech-Small-Molecule Drug Interactions 96%
- FLONE: fully Lorentz network embedding for inferring novel drug targets 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.