LLMsFold: Integrating Large Language Models and Biophysical Simulations for De Novo Drug Design
Waththe Liyanage, W. W.; Bove, F.; Righelli, D.; Romano, S.; Visone, R.; Iorio, M. V.; Lio, P.; Taccioli, C.
Show abstract
The discovery of novel small molecules is challenging because of the vastness of chemical space and the complexity of protein-ligand interactions, leading to low success rates and time-consuming workflows. Here, we present LLMsFold, a computational framework that combines Large Language Models (LLMs) and biophysical foundation tools to design and validate new small molecules targeting pathogenic proteins. The pipeline starts by identifying viable binding pockets on a target protein through geometry-based pocket detection. A 70-billion-parameter transformer model from the LlaMA family then generates candidate molecules as SMILES strings under prompt constraints that enforce drug-likeness. Each molecule is evaluated by Boltz-2, a diffusion-based model for protein-ligand co-folding that predicts bound 3D structure and binding affinity. Promising candidates are iteratively optimized through a reinforcement learning loop that prioritizes high predicted affinity and synthetic accessibility. We demonstrate the approach on two challenging targets: ACVR1 (Activin A Receptor Type 1), implicated in fibrodysplasia ossificans progressiva (FOP), and CD19, a surface antigen expressed on most B-cell lymphoma and leukemia cells. Top candidates show strong in silico binding predictions and favorable drug-like profiles. All code and models are made available to support reproducibility and further development.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The use of a graph database is a complementary approach to a classical similarity search for identifying commercially available fragment merges 97%
- Retro Drug Design: From Target Properties to Molecular Structures 96%
- Streamlining Computational Fragment-Based Drug Discovery through Evolutionary Optimization Informed by Ligand-Based Virtual Prescreening 95%
Similar papers in this journal
- NRGSuite-Qt: A PyMOL plugin for high-throughput virtual screening, molecular docking, normal-mode analysis, the study of molecular interactions and the detection of binding-site similarities 94%
- A Random Forest Classifier for Protein-Protein Docking Models 93%
- Improving classification of correct and incorrect protein-protein docking models by augmenting the training set 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.