Back

Biomedical Large Language Models and Prompt Engineering for Causality Assessment of Individual Case Safety Reports in Pharmacovigilance

Heckmann, N. S.; Papoutsi, D. G.; Barbieri, M. A.; Battini, V.; Molgaard, S. N.; Schmidt, S. O.; Melskens, L.; Sessa, M.

2026-02-24 pharmacology and therapeutics
10.64898/2026.02.19.26346467 medRxiv
Show abstract

BackgroundBiomedical Large Language Models (LLMs) combined with prompt engineering offer domain-specific reasoning, yet their application to individual-level causality assessment remains unexplored. This study evaluated five combinations of biomedical LLMs, prompting strategies, and causality algorithms by comparing their agreement with two human expert evaluators. Research design and methodsA total of 150 Individual Case Safety Reports (ICSRs) were analyzed: 140 reports from Food and Drug Administration Adverse Event Reporting System (FAERS), and 10 myocarditis/pericarditis ICSRs from Vaccine AERS (VAERS). Assessments were conducted using the Naranjo and WHO-UMC algorithms. Biomedical LLMs tested included TinyLlama 1.1B, Medicine LLaMA-3 8B, and MedLLaMA v20, combined with Chain-of-Thought (CoT) or Decomposition prompting. Agreement was measured using Gwets Agreement Coefficient 1 (AC1) and percentage agreement, alongside performance metrics and qualitative error analysis. ResultsThe Medicine LLaMA-3 8B-Naranjo-CoT combination achieved the highest agreement with human assessors for the final classification of causality (64%). Biomedical LLMs demonstrated low inter-rater agreement on critical items of causality assessment such as identification of listed AE, temporal plausibility, alternative causes, and objective evidence of AEs. Frequent model failures included irrelevant responses. ConclusionsBiomedical LLMs showed improved performance over general purpose models previously tested but remain suboptimal for reliable causality assessment of ICSRs.

Published in Pharmaceutical Research · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Drug Safety
10 papers in training set
Top 0.1%
9.7%
2
BioData Mining
22 papers in training set
Top 0.1%
7.8%
3
Clinical Pharmacology & Therapeutics
25 papers in training set
Top 0.1%
7.8%
4
Frontiers in Pharmacology
111 papers in training set
Top 0.2%
7.2%
5
Pharmacoepidemiology and Drug Safety
18 papers in training set
Top 0.1%
5.5%
6
Journal of Medical Internet Research
87 papers in training set
Top 0.4%
5.4%
7
PLOS ONE
5266 papers in training set
Top 30%
5.1%
8
Journal of Biomedical Informatics
47 papers in training set
Top 0.3%
4.8%
50% of probability mass above
9
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.7%
4.8%
10
npj Digital Medicine
118 papers in training set
Top 1%
3.5%
11
Systematic Reviews
15 papers in training set
Top 0.1%
3.2%
12
Clinical and Translational Science
22 papers in training set
Top 0.1%
3.2%
13
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
3.1%
14
British Journal of Clinical Pharmacology
21 papers in training set
Top 0.2%
2.1%
15
Scientific Reports
3612 papers in training set
Top 59%
1.5%
16
Clinical Trials
11 papers in training set
Top 0.2%
1.4%
17
JAMIA Open
42 papers in training set
Top 1%
1.3%
18
BMC Medical Research Methodology
47 papers in training set
Top 0.9%
1.3%
19
Schizophrenia
21 papers in training set
Top 0.3%
1.1%
20
Value in Health
11 papers in training set
Top 0.3%
1.1%
21
Journal of Allergy and Clinical Immunology
27 papers in training set
Top 0.4%
1.0%
22
Epilepsia
56 papers in training set
Top 0.5%
1.0%
23
Frontiers in Medicine
120 papers in training set
Top 4%
0.8%
24
BJGP Open
13 papers in training set
Top 0.5%
0.8%
25
JAMA Network Open
130 papers in training set
Top 5%
0.6%
26
Computers in Biology and Medicine
128 papers in training set
Top 5%
0.6%
27
Communications Medicine
113 papers in training set
Top 6%
0.6%
28
The American Journal of Tropical Medicine and Hygiene
68 papers in training set
Top 2%
0.6%
29
eBioMedicine
183 papers in training set
Top 8%
0.6%