Back

Chemical Probes in Scientific Literature: Expanding and Validating Target-Disease Evidence

Adasme, M. F.; Ochoa, D.; Lopez, I.; Do, H.-M.-A.; McDonagh, E. M.; O'Boyle, N. M.; Leach, A. R.; Zdrazil, B.

2026-02-20 bioinformatics
10.64898/2026.02.19.706919 bioRxiv
Show abstract

Chemical probes are indispensable tools for validating therapeutic hypotheses, yet their broader impact on early-stage drug discovery remains unquantified. To our knowledge, this study represents the first systematic, large-scale investigation of the chemical probe literature. By screening over 18 million articles using a high-quality dictionary of 561 chemical probes, we identified 20,000 articles mentioning a chemical probe which resulted in 5,558 unique target-disease (T-D) associations. Our analysis yields four principal findings that redefine the utility of these chemicals: First, we show that chemical probe evidence typically precedes the appearance of structured data in major knowledge bases by 1-7 years, providing a crucial lead time for target prioritisation. Second, we identified 353 T-D pairs (6.4%) with no prior evidence in the Open Targets Platform, highlighting the approachs discovery potential. Third, the application of strict novelty filters uncovered 135 new high-confidence associations between targets and diseases, revealing distinct opportunities for therapeutic repurposing in non-oncological, rare autoimmune diseases, and diseases without effective therapies due to complex biology or high treatment resistance. Finally, we demonstrate that chemical probes are essential for strengthening evidence, providing functional validation for associations previously supported only by weaker, correlative data such as RNA expression or animal models. Collectively, these findings illustrate that chemical probes catalyse early therapeutic discovery, emphasising the importance of cataloguing existing probes and identifying new ones.

Matching journals

The top 16 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.