Back

Clinical Advancement Forecasting

Czech, E. A.; Wojdyla, R. S.; Himmelstein, D. S.; Frank, D. H.; Miller, N. A.; Milwid, J. M.; Kolom, A.; Hammerbacher, J.

2024-08-03 genetic and genomic medicine
10.1101/2024.08.02.24311422 medRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWChoosing which drug targets to pursue for a given disease is one of the most impactful decisions made in the global development of new medicines. This study examines the extent to which the outcomes of clinical trials can be predicted based on a small set of longitudinal (temporally labeled) evidence and properties of drug targets and diseases. We demonstrate a novel statistical learning framework for identifying the top 2% of target-disease pairs that are as much as 4-5x more likely to advance beyond phase 2 trials. This framework is 1.5-2x more effective than an Open Targets composite score based on the same set of evidence. It is also 2x more effective than a common measure for genetic support that has been observed previously, as well as in this study, to confer a 2x higher likelihood of success. Utilizing a subset of our biomedical evidence base, non-negative linear models resulting from this framework can produce simple weighting schemes across various types of human, animal, and cell model genomic, transcriptomic, proteomic, and clinical evidence to identify previously undeveloped target-disease pairs poised for clinical success. In this study we further explore: i) how longitudinal treatment of evidence relates to leakage and reverse causality in biomedical research and how temporalized evidence can mitigate common forms of potential biases and inflation ii) the relative impact of different types of features on our predictions; and iii) an analysis of the space of currently undeveloped, tractable targets predicted with these methods to have the highest likelihood of clinical success. To ease reproduction and deployment, no data is used outside of Open Targets and the described methods require no expert knowledge, and can support expansion of lines of evidence to further improve performance.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 3%
9.3%
2
Communications Medicine
113 papers in training set
Top 0.2%
7.0%
3
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.1%
6.0%
4
npj Digital Medicine
118 papers in training set
Top 1%
5.3%
5
PLOS ONE
5266 papers in training set
Top 30%
5.0%
6
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.1%
5.0%
7
PLOS Computational Biology
1863 papers in training set
Top 8%
4.7%
8
Scientific Reports
3612 papers in training set
Top 21%
4.7%
9
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.8%
4.2%
50% of probability mass above
10
Genome Medicine
183 papers in training set
Top 1%
3.4%
11
International Journal of Molecular Sciences
494 papers in training set
Top 4%
3.1%
12
Journal of Biomedical Informatics
47 papers in training set
Top 0.5%
3.1%
13
Nature Communications
5641 papers in training set
Top 37%
3.1%
14
Frontiers in Bioinformatics
49 papers in training set
Top 0.2%
2.7%
15
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 20%
2.7%
16
iScience
1154 papers in training set
Top 10%
2.5%
17
eLife
5828 papers in training set
Top 43%
2.3%
18
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.0%
19
Patterns
78 papers in training set
Top 1%
1.7%
20
Genome Biology
637 papers in training set
Top 7%
1.1%
21
Cell Genomics
172 papers in training set
Top 3%
1.1%
22
The American Journal of Human Genetics
234 papers in training set
Top 2%
1.1%
23
PLOS Genetics
862 papers in training set
Top 11%
1.0%
24
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.0%
25
Frontiers in Genetics
230 papers in training set
Top 6%
0.8%
26
Human Genetics and Genomics Advances
84 papers in training set
Top 2%
0.8%
27
Bioinformatics Advances
203 papers in training set
Top 5%
0.8%
28
Molecular Systems Biology
162 papers in training set
Top 3%
0.8%
29
GigaScience
212 papers in training set
Top 6%
0.6%
30
Nature Biomedical Engineering
47 papers in training set
Top 2%
0.6%