Diagnostic Accuracy of a locally deployed Large Language Model Algorithm for Automated Code Stroke Pathway Identification in Emergency Department Triage Notes
Valente, M. J.; Sharobeam, A.; Vuong, J.; Chan, W. K.
Show abstract
Background: Delayed Code Stroke activation contributes to worse outcomes in acute stroke. Emergency Department (ED) triage notes contain free-text clinical information that could enable automated, real-time pathway activation. We evaluated the diagnostic accuracy of a multi-pass large language model (LLM) pipeline for identifying patients meeting Code Stroke criteria from ED triage notes. Methods: A retrospective cross-sectional study was conducted at Monash Medical Centre, Melbourne, Australia. De-identified triage notes from 3,023 ED presentations over a one-month period (September-October 2023) were analysed. The pipeline applied sequential passes for translation, stroke symptom identification, mimic exclusion, baseline functional status, temporal window classification, and symptom resolution. Six locally deployed language models were evaluated. Performance was assessed against two reference standards: neurologist-labelled diagnosis and documented ED Code Stroke activation. Primary outcomes were sensitivity and specificity; secondary outcomes included PPV, NPV, and Gwet's AC1. Reliability of the neurologist reference standard was assessed by blinded independent re-review of a stratified random sample of 200 presentations by a second neurologist. Results: Of 3,023 presentations, 136 were neurologist-labelled positive. Agreement between the primary and a blinded second neurologist on a 200-note reliability sub-sample was almost perfect (raw agreement 95.0%, Cohen's K; 0.900, 95% CI 0.838-0.959). The cohort included 140 ED Code Stroke activations (median age 69, IQR 56-81 years), of whom 83 (59.2%) had confirmed stroke diagnosis. Sixteen patients (11.4%) underwent endovascular clot retrieval and 4 (2.9%) received thrombolysis. The best-performing model (Qwen 2.5 14B) achieved sensitivity 0.890 (95% CI 0.826-0.932), specificity 0.993 (0.989-0.996), PPV 0.858 (0.791-0.906), and NPV 0.995 (0.991-0.997). Pairwise McNemar testing demonstrated statistically superior overall accuracy for Qwen 2.5 14B over Llama 3.1 8B, Phi-4 14B, and Mistral 3 14B (all p<0.001 after Holm correction), with no significant difference detected versus Nemotron-Nano-12B-v2 or Qwen 3 14B. Conclusions: A locally deployed language model demonstrates acceptable sensitivity and specificity for automated Code Stroke identification from free-text triage notes. Performance was comparable across the two best models, suggesting that capable open-weight models in this parameter range may be sufficient to proceed with ongoing internal testing and external validation. The pipeline operates without internet connectivity or model retraining on patient data, supporting feasibility for real-world ED integration.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 90%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 89%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 89%
Similar papers in this journal
- Automated Identification of Thrombectomy Amenable Vessel Occlusion on Computed Tomography Angiography using Deep Learning 90%
- Workflow Intervals andOutcomesof Endovascular Treatment for Acute Large-Vessel Occlusion During On- Versus Off-Hours in China The ANGEL-ACT Registry 89%
- Factors Associated with Stroke after COVID-19 Vaccination: A Statewide Analysis 89%
Similar papers in this journal
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 95%
- Imputation strategies for missing baseline neurological assessment covariates after traumatic brain injury: A CENTER-TBI study 92%
- Leveraging Machine Learning for Enhanced and Interpretable Risk Prediction of Venous Thromboembolism in Acute Ischemic Stroke Care 92%
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 90%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 90%
- Accuracy of preferred language data in a multi-hospital electronic health record in Toronto, Canada 90%
Similar papers in this journal
- A hybrid simulation-based workshop improves knowledge and confidence in the management of hemorrhagic conversion of stroke among interventional neurology trainees 92%
- Prehospital triage of intracranial hemorrhage and anterior large vessel occlusion ischemic stroke: the value of the rapid arterial occlusion evalution 92%
- Large Core Thrombectomy: Feasibility Of Simplified Protocol In Resource-Limited Settings 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.