Clinical Advancement Forecasting
Czech, E. A.; Wojdyla, R. S.; Himmelstein, D. S.; Frank, D. H.; Miller, N. A.; Milwid, J. M.; Kolom, A.; Hammerbacher, J.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWChoosing which drug targets to pursue for a given disease is one of the most impactful decisions made in the global development of new medicines. This study examines the extent to which the outcomes of clinical trials can be predicted based on a small set of longitudinal (temporally labeled) evidence and properties of drug targets and diseases. We demonstrate a novel statistical learning framework for identifying the top 2% of target-disease pairs that are as much as 4-5x more likely to advance beyond phase 2 trials. This framework is 1.5-2x more effective than an Open Targets composite score based on the same set of evidence. It is also 2x more effective than a common measure for genetic support that has been observed previously, as well as in this study, to confer a 2x higher likelihood of success. Utilizing a subset of our biomedical evidence base, non-negative linear models resulting from this framework can produce simple weighting schemes across various types of human, animal, and cell model genomic, transcriptomic, proteomic, and clinical evidence to identify previously undeveloped target-disease pairs poised for clinical success. In this study we further explore: i) how longitudinal treatment of evidence relates to leakage and reverse causality in biomedical research and how temporalized evidence can mitigate common forms of potential biases and inflation ii) the relative impact of different types of features on our predictions; and iii) an analysis of the space of currently undeveloped, tractable targets predicted with these methods to have the highest likelihood of clinical success. To ease reproduction and deployment, no data is used outside of Open Targets and the described methods require no expert knowledge, and can support expansion of lines of evidence to further improve performance.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TUGDA: Task uncertainty guided domain adaptation for robust generalization of cancer drug response prediction from in vitro to in vivo settings 94%
- Looking at the BiG picture: Incorporating bipartite graphs in drug response prediction 94%
- MOViDA: Multi-Omics Visible Drug Activity Prediction with a Biologically Informed Neural Network Model 94%
Similar papers in this journal
- Drug-combination wide association studies of cancer 92%
- Bayesian combination of mechanistic modeling and machine learning (BaM3): improving personalized tumor growth predictions 92%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 92%
Similar papers in this journal
- piCRISPR: Physically Informed Deep Learning Models for CRISPR/Cas9 Off-Target Cleavage Prediction 91%
- Federated Learning for Predicting Compound Mechanism of Action Based on Image-data from Cell Painting 91%
- Unraveling the Co-Morbidity between COVID-19 and Neurodegenerative Diseases Through Multi-scale Graph Analysis: A Systematic Investigation of Biological Databases and Text Mining 90%
Similar papers in this journal
Similar papers in this journal
- MatchMaker: A Deep Learning Framework for Drug Synergy Prediction 94%
- Genetic analysis of coronary artery disease using tree-based automated machine learning informed by biology-based feature selection 93%
- Latent representation of the human pan-celltype epigenome through a deep recurrent neural network 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.