Back

Cross-Model Variability in Large Language Model Triage Behavior for Potential Stroke Symptoms

Dworkis, D. A.; Stenstrom, J.; Sen, A.; Lucarelli, R. T.

2026-05-25 emergency medicine
10.64898/2026.05.22.26353904 medRxiv
Show abstract

Background: Stroke is a time-sensitive neurological emergency in which early EMS activation and presentation to definitive care are cornerstones of effective therapy. Large language models (LLMs) are increasingly consulted by the public for medical advice, but the veracity of the guidance provided by commercially available models responding to potential stroke symptoms is not well understood. Methods: We performed a cross-model benchmarking study comparing the triage choices of three frontier LLMs (Claude Sonnet 4.6, GPT-4o, and Llama 3.3-70b-versatile) on first-person vignettes describing a unilateral arm symptom on waking, across 10 symptom descriptors, and two clinical phases (before and after a partially reassuring self-examination), with or without a clinical distractor (n=50 per condition). Results: Claude sought emergency care most often, Llama least, and GPT-4o in between, diverging most sharply in the post-examination phase where Claude called 911 in 100% of runs, Llama called for non-emergency help in 100%, and GPT-4o was symptom-dependent. A distractor shifted behavior away from emergency care in almost all conditions: calling 911 fell from 37.9% to 14.6% and waiting rose from 0% to 45.9% in the post-examination vignette. Responses were also sensitive to symptom word: weak, limp, heavy, and clumsy generated higher alarm, whereas numb, tingly, odd, strange, and weird generated less urgent responses. Conclusions: The increasing use of LLMs for medical advice has significant public health implications. Commercially available LLMs show significant model-to-model variability and framing sensitivity when confronted with potential stroke symptoms, including under-recognition of canonical CDC warning descriptors, underscoring the need for systematic benchmarking as these tools become de facto first points of contact for patients experiencing neurological emergencies.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Stroke
40 papers in training set
Top 0.1%
15.8%
2
PLOS ONE
5266 papers in training set
Top 15%
12.5%
3
Emergency Medicine Journal
21 papers in training set
Top 0.1%
10.2%
4
Scientific Reports
3612 papers in training set
Top 9%
7.0%
5
CMAJ Open
12 papers in training set
Top 0.1%
5.1%
50% of probability mass above
6
BMC Health Services Research
51 papers in training set
Top 0.4%
5.1%
7
JAMA Network Open
130 papers in training set
Top 0.9%
3.7%
8
Artificial Intelligence in Medicine
17 papers in training set
Top 0.1%
3.4%
9
Journal of the American Heart Association
140 papers in training set
Top 2%
2.6%
10
npj Digital Medicine
118 papers in training set
Top 2%
2.5%
11
PLOS Digital Health
106 papers in training set
Top 2%
2.1%
12
PLOS Global Public Health
344 papers in training set
Top 5%
2.0%
13
Heliyon
152 papers in training set
Top 4%
1.4%
14
Cureus
68 papers in training set
Top 3%
1.4%
15
Communications Medicine
113 papers in training set
Top 3%
1.2%
16
Journal of Stroke and Cerebrovascular Diseases
15 papers in training set
Top 0.4%
1.2%
17
BioData Mining
22 papers in training set
Top 0.4%
1.2%
18
Frontiers in Public Health
148 papers in training set
Top 4%
1.2%
19
Archives of Physical Medicine and Rehabilitation
10 papers in training set
Top 0.3%
1.1%
20
Medicine
31 papers in training set
Top 2%
0.9%
21
Frontiers in Neurology
102 papers in training set
Top 3%
0.6%
22
Alzheimer's & Dementia: Translational Research & Clinical Interventions
17 papers in training set
Top 0.6%
0.6%
23
Healthcare
17 papers in training set
Top 1%
0.6%
24
Frontiers in Medicine
120 papers in training set
Top 5%
0.6%
25
BMJ Open
601 papers in training set
Top 14%
0.5%
26
Journal of The Royal Society Interface
235 papers in training set
Top 5%
0.5%