Lessons learned during the journey of data: from experiment to model for predicting kinase affinity, selectivity, polypharmacology, and resistance
Lopez-Rios de Castro, R.; Rodriguez-Guerra, J.; Schaller, D.; Kimber, T. B.; Taylor, C.; White, J. B.; Backenkohler, M.; Payne, A.; Kaminow, B.; Pulido, I.; Singh, S.; Krammer, P. L.; Perez-Hernandez, G.; Volkamer, A.; Chodera, J. D.
Show abstract
Recent advances in machine learning (ML) are reshaping drug discovery. Structure-based ML methods use physically-inspired models to predict binding affinities from protein:ligand complexes. These methods promise to enable the integration of data for many related targets, which addresses issues related to data scarcity for single targets and could enable generalizable predictions for a broad range of targets, including mutants. In this work, we report our experiences in building KinoML, a novel framework for ML in target-based small molecule drug discovery with an emphasis on structure-enabled methods. KinoML focuses currently on kinases as the relative structural conservation of this protein superfamily, particularly in the kinase domain, means it is possible to leverage data from the entire superfamily to make structure-informed predictions about binding affinities, selectivities, and drug resistance. Some key lessons learned in building KinoML include: the importance of reproducible data collection and deposition, the harmonization of molecular data and featurization, and the choice of the right data format to ensure reusability and reproducibility of ML models. As a result, KinoML allows users to easily achieve three tasks: accessing and curating molecular data; featurizing this data with representations suitable for ML applications; and running reproducible ML experiments that require access to ligand, protein, and assay information to predict ligand affinity. Despite KinoML focusing on kinases, this framework can be applied to other proteins. The lessons reported here can help guide the development of platforms for structure-enabled ML in other areas of drug discovery.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NRGSuite-Qt: A PyMOL plugin for high-throughput virtual screening, molecular docking, normal-mode analysis, the study of molecular interactions and the detection of binding-site similarities 95%
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 94%
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 94%
Similar papers in this journal
Similar papers in this journal
- mdciao: Accessible Analysis and Visualization of Molecular Dynamics Simulation Data 95%
- Protein Domain-Based Prediction of Compound-Target Interactions and Experimental Validation on LIM Kinases 94%
- Towards a comprehensive view of the pocketome universe - biological implications and algorithmic challenges. 94%
Similar papers in this journal
- From Library to Landscape: Integrative Annotation Workflows for Compound Libraries in Drug Repurposing 94%
- SynLethDB 2.0: A web-based knowledge graph database on synthetic lethality for novel anticancer drug discovery 92%
- Peptipedia v2.0: A peptide sequence database and user-friendly web platform. A major update 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.