GatorAffinity: Boosting Protein-Ligand Binding Affinity Prediction with Large-Scale Synthetic Structural Data
Wei, J.; Zhang, Y.; Ramdhan, P. A.; Huang, Z.; Seabra, G.; Jiang, Z.; Li, C.; Li, Y.
Show abstract
Protein-ligand binding affinity prediction is a fundamental task in computational drug discovery. Although substantial efforts have been made to enhance prediction accuracy using data-driven approaches, progress remains limited by persistent data scarcity. The widely used PDBbind dataset, for example, contains fewer than 20, 000 experimental structures with annotated binding affinities, while a vast number of affinity measurements remain underutilized due to missing structural data. Here, we investigate this untapped potential by curating more than 450, 000 synthetic protein-ligand complexes annotated with Kd and Ki values using the Boltz-1 structure prediction model. Building on this unprecedented scale of synthetic data, further augmented with over 1 million synthetic complexes from the recently released SAIR database annotated with IC50 values, we develop GatorAffinity, a geometric deep learning-based scoring function pretrained on large-scale synthetic data and fine-tuned using high-quality experimental structures from PDBbind. Extensive evaluation on a leak-proof benchmark demonstrates that GatorAffinity significantly outperforms state-of-the-art affinity prediction methods, offering superior accuracy and generalizability. Our findings show that augmenting available experimental data with synthetic complexes can effectively address the data scarcity challenge while maintaining strong predictive reliability. By releasing the pretrained GatorAffinity model and the large-scale synthetic dataset GatorAffinity-DB, we provide a scalable and reproducible foundation for affinity prediction, virtual screening, and broader structure-based drug design applications (https://github.com/AIDD-LiLab/GatorAffinity).
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From Proteins to Ligands: Decoding Deep Learning Methods for Binding Affinity Prediction 97%
- Accurate Conformation Sampling via Protein Structural Diffusion 97%
- ProAffinity-GNN: A Novel Approach to Structure-based Protein-Protein Binding Affinity Prediction via a Curated Dataset and Graph Neural Networks 96%
Similar papers in this journal
- To pack or not to pack: revisiting protein side-chain packing in the post-AlphaFold era 96%
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 96%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 95%
Similar papers in this journal
- Ig-VAE: Generative Modeling of Immunoglobulin Proteins by Direct 3D Coordinate Generation 96%
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 96%
- Hybridized distance- and contact-based hierarchical structure modeling for folding soluble and membrane proteins 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.