Large language model aided automatic high-throughput drug screening using self-controlled cohort study
Xu, S.; Finkelstein, S. N.; Welsch, R. E.; Ng, K.; Tzoulaki, I.; Middleton, L.
Show abstract
BackgroundDeveloping medicine from scratch to governmental authorization and detecting adverse drug reactions (ADR) have barely been economical, expeditious, and risk-averse investments. The availability of large-scale observational healthcare databases and the popularity of large language models offer an unparalleled opportunity to enable automatic high-throughput drug screening for both repurposing and pharmacovigilance. ObjectivesTo demonstrate a general workflow for automatic high-throughput drug screening with the following advantages: (i) the association of various exposure on diseases can be estimated; (ii) both repurposing and pharmacovigilance are integrated; (iii) accurate exposure length for each prescription is parsed from clinical texts; (iv) intrinsic relationship between drugs and diseases are removed jointly by bioinformatic mapping and large language model - ChatGPT; (v) causal-wise interpretations for incidence rate contrasts are provided. MethodsUsing a self-controlled cohort study design where subjects serve as their own control group, we tested the intention-to-treat association between medications on the incidence of diseases. Exposure length for each prescription is determined by parsing common dosages in English free text into a structured format. Exposure period starts from initial prescription to treatment discontinuation. A same exposure length preceding initial treatment is the control period. Clinical outcomes and categories are identified using existing phenotyping algorithms. Incident rate ratios (IRR) are tested using uniformly most powerful (UMP) unbiased tests. ResultsWe assessed 3,444 medications on 276 diseases on 6,613,198 patients from the Clinical Practice Research Datalink (CPRD), an UK primary care electronic health records (EHR) spanning from 1987 to 2018. Due to the built-in selection bias of self-controlled cohort studies, ingredients-disease pairs confounded by deterministic medical relationships are removed by existing map from RxNorm and nonexistent maps by calling ChatGPT. A total of 16,901 drug-disease pairs reveals significant risk reduction, which can be considered as candidates for repurposing, while a total of 11,089 pairs showed significant risk increase, where drug safety might be of a concern instead. ConclusionsThis work developed a data-driven, nonparametric, hypothesis generating, and automatic high-throughput workflow, which reveals the potential of natural language processing in pharmacoepidemiology. We demonstrate the paradigm to a large observational health dataset to help discover potential novel therapies and adverse drug effects. The framework of this study can be extended to other observational medical databases.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 92%
- Using Natural Language Processing of Clinical Notes to Supplement Structured Electronic Health Record Data for Phenotyping Smoking and Obesity in a Healthcare System 92%
- Bias amplification of unobserved confounding in pharmacoepidemiological studies using indication-based sampling: there is no free lunch in restricting the sample to those with a particular drug-indication 91%
Similar papers in this journal
- Clinically Informed Semi-Supervised Learning Improves Disease Annotation and Equity from Electronic Health Records: A Glaucoma Case Study 92%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 91%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 91%
Similar papers in this journal
- Large-scale empirical identification of candidate comparators for pharmacoepidemiological studies 94%
- Standardization of drug names in the FDA Adverse Event Reporting System: The DiAna dictionary 92%
- Patient-Reported Reasons for Antihypertensive Medication Change: A Quantitative Study Using Social Media 90%
Similar papers in this journal
- Framework for Identifying Drug Repurposing Candidates from Observational Healthcare Data 95%
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 94%
- Using indication embeddings to represent patient health for drug safety studies 93%
Similar papers in this journal
- Causal reasoning over knowledge graphs leveraging drug-perturbed and disease-specific transcriptomic signatures for drug discovery 93%
- Benchmarking network propagation methods for disease gene identification 91%
- OpenABM-Covid19 - an agent-based model for non-pharmaceutical interventions against COVID-19 including contact tracing 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.