Deep Learning for Virtual Screening: Five Reasons to Use ROC Cost Functions
Golkov, V.; Becker, A.; Plop, D. T.; Cuturilo, D.; Davoudi, N.; Mendenhall, J.; Moretti, R.; Meiler, J.; Cremers, D.
Show abstract
Computer-aided drug discovery is an essential component of modern drug development. Therein, deep learning has become an important tool for rapid screening of billions of molecules in silico for potential hits containing desired chemical features. Despite its importance, substantial challenges persist in training these models, such as severe class imbalance, high decision thresholds, and lack of ground truth labels in some datasets. In this work we argue in favor of directly optimizing the receiver operating characteristic (ROC) in such cases, due to its robustness to class imbalance, its ability to compromise over different decision thresholds, certain freedom to influence the relative weights in this compromise, fidelity to typical benchmarking measures, and equivalence to positive/unlabeled learning. We also propose new training schemes (coherent mini-batch arrangement, and usage of out-of-batch samples) for cost functions based on the ROC, as well as a cost function based on the logAUC metric that facilitates early enrichment (i.e. improves performance at high decision thresholds, as often desired when synthesizing predicted hit compounds). We demonstrate that these approaches outperform standard deep learning approaches on a series of PubChem high-throughput screening datasets that represent realistic and diverse drug discovery campaigns on major drug target families.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AITL: Adversarial Inductive Transfer Learning with input and output space adaptation for pharmacogenomics 95%
- DTI-Voodoo: machine learning over interaction networks and ontology-based background knowledge predicts drug-target interactions 95%
- Neural Collective Matrix Factorization for Integrated Analysis of Heterogeneous Biomedical Data 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Towards explainable interaction prediction: Embedding biological hierarchies into hyperbolic interaction space 95%
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 94%
- A cautionary tale about properly vetting datasets used in supervised learning predicting metabolic pathway involvement 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.