Step Forward Cross Validation for Bioactivity Prediction: Out of Distribution Validation in Drug Discovery
Saha, U. S.; Vendruscolo, M.; Carpenter, A. E.; Singh, S.; Bender, A.; Seal, S.
Show abstract
Recent advances in machine learning methods for materials science have significantly enhanced accurate predictions of the properties of novel materials. Here, we explore whether these advances can be adapted to drug discovery by addressing the problem of prospective validation - the assessment of the performance of a method on out-of-distribution data. First, we tested whether k-fold n-step forward cross-validation could improve the accuracy of out-of-distribution small molecule bioactivity predictions. We found that it is more helpful than conventional random split cross-validation in describing the accuracy of a model in real-world drug discovery settings. We also analyzed discovery yield and novelty error, finding that these two metrics provide an understanding of the applicability domain of models and an assessment of their ability to predict molecules with desirable bioactivity compared to other small molecules. Based on these results, we recommend incorporating a k-fold n-step forward cross-validation and these metrics when building state-of-the-art models for bioactivity prediction in drug discovery.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data 96%
- Comprehensive machine learning boosts structure-based virtual screening for PARP1 inhibitors 94%
- qHTSWaterfall: 3-dimensional visualization software for quantitative high-throughput screening (qHTS) data 94%
Similar papers in this journal
Similar papers in this journal
- Machine learning approaches identify chemical features for stage-specific antimalarial compounds 96%
- How far are we in the rapid prediction of drug resistance caused by kinase mutations? 93%
- Utilizing Heteroatom Types and Numbers from Extensive Ligand Libraries to Develop Novel hERG Blocker QSAR Models Using Machine Learning-based Classifiers 93%
Similar papers in this journal
- Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo) 96%
- Intense bitterness of molecules: machine learning for expediting drug discovery 93%
- G-PLIP: Knowledge graph neural network for structure-free protein-ligand bioactivity prediction 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.