Back

Drug-drug interaction identification using large language models

Blotske, K.; Zhao, X.; Henry, K.; Gao, Y.; Tilley, A.; Cargile, M.; Murray, B.; Smith, S. E.; Barreto, E.; Bauer, S.; Sohn, S.; Liu, T.; Sikora, A.

2025-12-04 pharmacology and therapeutics
10.64898/2025.12.03.25341549 medRxiv
Show abstract

BackgroundDrug-drug interactions (DDIs) are a significant source of morbidity and adverse drug events (ADEs), particularly in situations of polypharmacy and complex medication regimens. While rules-based software integrated in electronic health records (EHRs) has demonstrated proficiency in identifying DDIs present in medication regimens, large language model (LLM) based identification requires thorough benchmarking and performance evaluation using high-quality datasets for safe use. The purpose of this study was to develop a series of performance benchmarking experiments specifically for LLM performance in identification and management of DDIs using a specifically curated clinician-annotated dataset of clinically-relevant DDIs. MethodsWe evaluated three LLMs (GPT-4o-mini, MedGemma-27B, LLaMA3-70B) using a clinician-annotated benchmark dataset of 750 DDI scenarios spanning three levels of diagnostic complexity. Tasks were aligned with flexible judgment formats: (1) a pointwise two-drug classification task, (2) a pairwise three-drug discrimination task, and (3) a listwise 4-6 drug selection task. Standardized zero-shot prompting with task-specific instructions was applied for all models. Performance was assessed using precision, recall, F1 score, and accuracy. Reliability was quantified using self-consistency across repeated runs and confidence-aligned metrics to capture stability in model reasoning. ResultsAcross the three experiments, model performance varied by task structure and interaction severity. LLaMA3-70B demonstrated the highest recall and F1 score in the pointwise task, whereas GPT-4o-mini achieved superior accuracy and consistency in the pairwise and listwise tasks. MedGemma-27B showed competitive performance in identifying Category D interactions. Self-consistency decreased as task complexity increased, highlighting reduced stability in multi-drug reasoning. No model exhibited uniformly high reliability across all judgment formats. ConclusionsCurrent LLMs show promising but uneven capabilities in identifying DDIs across clinically relevant task structures. Performance degrades as the reasoning space expands, and stability across repeated queries remains limited. These findings emphasize the need for multi-format evaluation frameworks and reliability-aware assessment when considering LLMs for medication-safety applications.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
BioData Mining
22 papers in training set
Top 0.1%
12.6%
2
Drug Safety
10 papers in training set
Top 0.1%
11.8%
3
Clinical Pharmacology & Therapeutics
25 papers in training set
Top 0.1%
7.8%
4
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.2%
7.8%
5
Frontiers in Pharmacology
111 papers in training set
Top 0.3%
6.2%
6
Journal of Medical Internet Research
87 papers in training set
Top 0.4%
5.1%
50% of probability mass above
7
JAMIA Open
42 papers in training set
Top 0.4%
4.3%
8
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.8%
4.3%
9
npj Digital Medicine
118 papers in training set
Top 1%
4.0%
10
PLOS ONE
5266 papers in training set
Top 38%
3.2%
11
British Journal of Clinical Pharmacology
21 papers in training set
Top 0.1%
3.2%
12
Journal of Biomedical Informatics
47 papers in training set
Top 0.5%
3.1%
13
Pharmacoepidemiology and Drug Safety
18 papers in training set
Top 0.1%
2.8%
14
Scientific Reports
3612 papers in training set
Top 54%
1.7%
15
Clinical and Translational Science
22 papers in training set
Top 0.3%
1.7%
16
JMIRx Med
32 papers in training set
Top 1%
1.3%
17
Schizophrenia
21 papers in training set
Top 0.3%
1.1%
18
Bioinformatics Advances
203 papers in training set
Top 4%
0.8%
19
International Journal of Medical Informatics
26 papers in training set
Top 1%
0.8%
20
Computers in Biology and Medicine
128 papers in training set
Top 5%
0.6%
21
eBioMedicine
183 papers in training set
Top 8%
0.6%
22
Journal of Allergy and Clinical Immunology
27 papers in training set
Top 0.6%
0.6%
23
Scientific Data
209 papers in training set
Top 3%
0.6%
24
Pharmacology Research & Perspectives
11 papers in training set
Top 0.4%
0.6%