High-Throughput Clinical Trial Emulation with Real World Data and Machine Learning: A Case Study of Drug Repurposing for Alzheimer's Disease
Zang, C.; Zhang, H.; Xu, J.; Zhang, H.; Fouladvand, S.; Havaldar, S.; Cheng, F.; Chen, K.; Chen, Y.; Glicksberg, B. S.; Chen, J.; Bian, J.; Wang, F.
Show abstract
Clinical trial emulation, which is the process of mimicking targeted randomized controlled trials (RCT) with real-world data (RWD), has attracted growing attention and interest in recent years from the pharmaceutical industry. Different from RCTs which have stringent eligibility criteria for recruiting participants, RWD are more representative of real-world patients to whom the drugs will be prescribed. One technical challenge for trial emulation is how to conduct effective confounding control with complex RWD so that the treatment effects can be objectively derived. Recently many approaches, including deep learning algorithms, have been proposed for this goal, but there is still no systematic evaluation and practical guidance on them. In this paper, we emulate 430, 000 trials from two large-scale RWD warehouses, covering both electronic health records (EHR) and general claims, over 170 million patients spanning more than 10 years, aiming to identify new indications of approved drugs for Alzheimers disease (AD). We have investigated the behaviors of multiple different approaches including logistic regression and deep learning models, and propose a new model selection strategy that can significantly improve the performance of confounding balance of the participants in different arms of emulated trials. We demonstrate that regularized logistic regression-based propensity score (PS) model outperforms the deep learning-based PS model and others, which contradicts with our intuitions to a certain extent. Finally, we identified 8 drugs whose original indications are not AD (pantoprazole, gabapentin, acetaminophen, atorvastatin, albuterol, fluticasone, amoxicillin, and omeprazole), hold great potential of being beneficial to AD patients.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 98%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
- Executable Network of SARS-CoV-2-Host Interaction Predicts Drug Combination Treatments 93%
Similar papers in this journal
- Machine learning guided association of adverse drug reactions with in vitro target-based pharmacology 95%
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 92%
- Multi-ancestry omic Mendelian randomization revealing putative drug targets of COVID-19 severity 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.