Back

Protocol for: A Simple, Accessible, Literature-based Drug Repurposing Pipeline

Lange, M.; Gogarty, E.; Martyn, M.; Braude, P.; Fayez, F.; Carter, B.

2024-07-19 health informatics
10.1101/2024.07.18.24310641 medRxiv
Show abstract

We will develop a novel approach to drug repurposing, utilising Natural Language Processing (NLP) and Literature Based Discovery (LBD) techniques. This will present a simplified, accessible drug repurposing pipeline using Word2Vec embeddings trained on PubMed abstracts to identify potential new medications to be repurposed. We present this approach in the context of antipsychotics, but it could be repeated for any available medication. The research is structured in three stages: O_LIIdentification of candidate medications using Word2Vec algorithm trained on scientific literature. C_LIO_LIEmpirical testing of identified candidates using a large hospital dataset to explore protective effects against disease onset. C_LIO_LIValidation of findings using a second, independent dataset to assess generalizability. C_LI This method addresses limitations in current machine learning-based drug repurposing approaches, including lack of external validation and limited accessibility. By leveraging Word2Vecs ability to capture semantic relationships between words, the study aims to uncover hidden connections in medical literature that may lead to novel therapeutic discoveries. The protocol emphasizes transparency and reproducibility, utilizing publicly available electronic health record (EHR) databases for validation. This approach allows for tangible results even for researchers with limited machine learning expertise, bridging the gap between biomedical and information systems communities.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.