Benchmarking Cross-Docking Strategies for Structure-Informed Machine Learning in Kinase Drug Discovery
Schaller, D.; Christ, C. D.; Chodera, J. D.; Volkamer, A.
Show abstract
In recent years machine learning has transformed many aspects of the drug discovery process including small molecule design for which the prediction of the bioactivity is an integral part. Leveraging structural information about the interactions between a small molecule and its protein target has great potential for downstream machine learning scoring approaches, but is fundamentally limited by the accuracy with which protein:ligand complex structures can be predicted in a reliable and automated fashion. With the goal of finding practical approaches to generating useful kinase:inhibitor complex geometries for downstream machine learning scoring approaches, we present a kinase-centric docking benchmark assessing the performance of different classes of docking and pose selection strategies to assess how well experimentally observed binding modes are recapitulated in a realistic crossdocking scenario. The assembled benchmark data set focuses on the well-studied protein kinase family and comprises a subset of 589 protein structures co-crystallized with 423 ATP-competitive ligands. We find that the docking methods biased by the co-crystallized ligand--utilizing shape overlap with or without maximum common substructure matching--are more successful in recovering binding poses than standard physics-based docking alone. Also, docking into multiple structures significantly increases the chance to generate a low RMSD docking pose. Docking utilizing an approach that combines all three methods (Posit) into structures with the most similar co-crystallized ligands according to shape and electrostatics proofed to be the most efficient way to reproduce binding poses achieving a success rate of 66.9 % across all included systems. The studied docking and pose selection strategies--which utilize the OpenEye Toolkit--were implemented into pipelines of the KinoML framework allowing automated and reliable protein:ligand complex generation for future downstream machine learning tasks. Although focused on protein kinases, we believe the general findings can also be transferred to other protein families.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Are Deep Learning Structural Models Sufficiently Accurate for Virtual Screening? Application of Docking Algorithms to AlphaFold2 Predicted Structures 98%
- ArtiDock: accurate Machine Learning approach to protein-ligand docking optimized for high-throughput virtual screening 96%
- Sampling and ranking of protein conformations using machine learning techniques do not improve quality of rigid protein-protein docking 95%
Similar papers in this journal
- RosettaGPCR: Multiple Template Homology Modeling of GPCRs with Rosetta 95%
- Protein Domain-Based Prediction of Compound-Target Interactions and Experimental Validation on LIM Kinases 94%
- Novel, provable algorithms for efficient ensemble-based computational protein design and their application to the redesign of the c-Raf-RBD:KRas protein-protein interface 94%
Similar papers in this journal
- Assessment of Software Methods for Estimating Protein-Protein Relative Binding Affinities 95%
- Deep learning based predictive modeling to screen natural compounds against TNF-alpha for the potential management of Rheumatoid Arthritis: Virtual screening to comprehensive in silico investigation 94%
- PharmaNet: Pharmaceutical discovery with deep recurrent neural networks. 94%
Similar papers in this journal
- Protein-protein docking with large-scale backbone flexibility using coarse-grained Monte-Carlo simulations 96%
- vScreenML v2.0: Improved Machine Learning Classification for Reducing False Positives in Structure-Based Virtual Screening 95%
- Single nucleotide polymorphism induces divergent dynamic patterns in CYP3A5: a microsecond scale biomolecular simulation of variants identified in Sub-Saharan African populations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.