SAIR: Enabling deep learning for protein-ligand lnteractions with a synthetic structural dataset
Lemos, P.; Beckwith, Z.; Bandi, S.; Van Damme, M.; Crivelli-Decker, J.; Shields, B. J.; Merth, T.; Jha, P. K.; De Mitri, N.; Callahan, T. J.; Nish, A.; Abruzzo, P.; Salomon-Ferrer, R.; Ganahl, M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWAccurate prediction of protein-ligand binding affinities remains a cornerstone problem in drug discovery. While binding affinity is inherently dictated by the 3D structure and dynamics of protein-ligand complexes, current deep learning approaches are limited by the lack of high-quality experimental structures with annotated binding affinities. To address this limitation, we introduce the Struc-turally Augmented IC50 Repository (SAIR), the largest publicly available dataset of protein-ligand 3D structures with associated activity data. The dataset com-prises 5, 244, 285 structures across 1, 048, 857 unique protein-ligand systems, cu-rated from the ChEMBL and BindingDB databases, which were then computa-tionally folded using the Boltz-1x model. We provide a comprehensive charac-terization of the dataset, including distributional statistics of proteins and ligands, and evaluate the structural fidelity of the folded complexes using PoseBusters. Our analysis reveals that approximately 3% of structures exhibit physical anoma-lies, predominantly related to internal energy violations. As an initial demon-stration, we benchmark several binding affinity prediction methods, including empirical scoring functions (Vina, Vinardo), a 3D convolutional neural network (Onionnet-2), and a graph neural network (AEV-PLIG). While machine learning-based models consistently outperform traditional scoring function methods, nei-ther exhibit a high correlation with ground truth affinities, highlighting the need for models specifically fine-tuned to synthetic structure distributions. This work provides a foundation for developing and evaluating next-generation structure and binding-affinity prediction models and offers insights into the structural and phys-ical underpinnings of protein-ligand interactions. The dataset can be found at https://www.sandboxaq.com/sair.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Robustly interrogating machine learning based scoring functions: what are they learning? 98%
- AHoJ: rapid, tailored search and retrieval of apo and holo protein structures for user-defined ligands 96%
- A Gated Graph Transformer for Protein ComplexStructure Quality Assessment and its Performancein CASP15 96%
Similar papers in this journal
- ArtiDock: accurate Machine Learning approach to protein-ligand docking optimized for high-throughput virtual screening 96%
- From Proteins to Ligands: Decoding Deep Learning Methods for Binding Affinity Prediction 95%
- RosENet: Improving binding affinity prediction by leveraging molecular mechanics energies with a 3D Convolutional Neural Network 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.