Back

Diagnostic Accuracy of a locally deployed Large Language Model Algorithm for Automated Code Stroke Pathway Identification in Emergency Department Triage Notes

Valente, M. J.; Sharobeam, A.; Vuong, J.; Chan, W. K.

2026-08-05 neurology
10.64898/2026.08.03.26359639 medRxiv
Show abstract

Background: Delayed Code Stroke activation contributes to worse outcomes in acute stroke. Emergency Department (ED) triage notes contain free-text clinical information that could enable automated, real-time pathway activation. We evaluated the diagnostic accuracy of a multi-pass large language model (LLM) pipeline for identifying patients meeting Code Stroke criteria from ED triage notes. Methods: A retrospective cross-sectional study was conducted at Monash Medical Centre, Melbourne, Australia. De-identified triage notes from 3,023 ED presentations over a one-month period (September-October 2023) were analysed. The pipeline applied sequential passes for translation, stroke symptom identification, mimic exclusion, baseline functional status, temporal window classification, and symptom resolution. Six locally deployed language models were evaluated. Performance was assessed against two reference standards: neurologist-labelled diagnosis and documented ED Code Stroke activation. Primary outcomes were sensitivity and specificity; secondary outcomes included PPV, NPV, and Gwet's AC1. Reliability of the neurologist reference standard was assessed by blinded independent re-review of a stratified random sample of 200 presentations by a second neurologist. Results: Of 3,023 presentations, 136 were neurologist-labelled positive. Agreement between the primary and a blinded second neurologist on a 200-note reliability sub-sample was almost perfect (raw agreement 95.0%, Cohen's K; 0.900, 95% CI 0.838-0.959). The cohort included 140 ED Code Stroke activations (median age 69, IQR 56-81 years), of whom 83 (59.2%) had confirmed stroke diagnosis. Sixteen patients (11.4%) underwent endovascular clot retrieval and 4 (2.9%) received thrombolysis. The best-performing model (Qwen 2.5 14B) achieved sensitivity 0.890 (95% CI 0.826-0.932), specificity 0.993 (0.989-0.996), PPV 0.858 (0.791-0.906), and NPV 0.995 (0.991-0.997). Pairwise McNemar testing demonstrated statistically superior overall accuracy for Qwen 2.5 14B over Llama 3.1 8B, Phi-4 14B, and Mistral 3 14B (all p<0.001 after Holm correction), with no significant difference detected versus Nemotron-Nano-12B-v2 or Qwen 3 14B. Conclusions: A locally deployed language model demonstrates acceptable sensitivity and specificity for automated Code Stroke identification from free-text triage notes. Performance was comparable across the two best models, suggesting that capable open-weight models in this parameter range may be sufficient to proceed with ongoing internal testing and external validation. The pipeline operates without internet connectivity or model retraining on patient data, supporting feasibility for real-world ED integration.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.