OHCA-EXTRACT: Evaluating the Accuracy of a Large Language Model Pipeline for Out-of-Hospital Cardiac Arrest Case Identification and Utstein Variable Extraction
Shekhar, A. C.; apakama, d.; Saharan, A.; Petrozzo, K.; Jiang, J.; Zebrowski, A.; Buckler, D. G.; Huebinger, R.; Abella, B.; Richardson, L. D.; Nadkarni, G.; Abbott, E.
Show abstract
Background: Manual identification and abstraction of out-of-hospital cardiac arrest (OHCA) cases and Utstein template variables from electronic health records is resource-intensive and limits scalable measurement for observational research and registry participation. We evaluated a large language model (LLM)-assisted pipeline to identify true OHCA encounters within an ICD-coded emergency department (ED) cohort and to also extract a limited set of Utstein variables from unstructured documentation, with iterative pipeline refinement and final performance evaluation conducted on the same validation cohort. Methods: We conducted a study of ICD-flagged OHCA encounters within a large urban academic health system (2015?2024). A two-module pipeline was developed that included an identification module and a variable-extraction module. The pipeline was evaluated against independent manual chart abstraction on a validation sample (n = 176) with physician adjudication. The identification module performed binary OHCA classification along with etiology classification. The variable-extraction module extracted five Utstein variables: witnessed status, EMS defibrillation, first recorded rhythm, arrest location, and bystander response. We calculated sensitivity, specificity, PPV, NPV, and F1 scores (Clopper?Pearson 95% CIs) using the manual chart abstraction as the reference gold standard. Results: Of 176 processed encounters, 152 had complete gold-standard classification and were included in identification analyses (OHCA prevalence 81.6%; n = 124 true OHCA, n = 28 non-OHCA). The identification module achieved an accuracy of 0.91 (95% CI 0.85?0.95), sensitivity 0.94 (0.89?0.98), specificity 0.75 (0.55?0.89), PPV 0.94 (0.89?0.98), NPV 0.75 (0.55?0.89), and F1 0.94, compared with accuracy 0.82 for a prevalence-only baseline. Variable-level accuracy ranged from 0.63 to 0.75 (Kappa range 0.31?0.51) across the five extracted Utstein variables, with bystander response showing the lowest agreement (accuracy 0.63, Kappa = 0.31). Conclusions: Within-sample performance estimates indicate that this LLM pipeline can identify true OHCA encounters within an ICD-flagged ED cohort with accuracy and PPV above an administrative-coding baseline. Variable-level extraction accuracy ranged from fair to moderate across the five evaluated Utstein variables, indicating that further methodological development is required before automated extraction is feasible. With external validation and integration with targeted human verification, this approach may support human-augmented workflows using LLMs, however independent external validation is required before our findings can be generalized.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 95%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
Similar papers in this journal
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
- Identification of physiological adverse events using continuous vital signs monitoring during paediatric critical care transport: a novel data-driven approach 92%
- Use of a Continuous Single Lead Electrocardiogram Analytic to Predict Patient Deterioration Requiring Rapid Response Team Activation 92%
Similar papers in this journal
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 94%
- Clinical prediction rule for SARS-CoV-2 infection from 116 U.S. emergency departments 93%
- Derivation and validation of a triage tool for acutely ill adults with suspected COVID-19: The PRIEST observational cohort study 93%
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 90%
- GenECG: A synthetic image-based ECG dataset to augment artificial intelligence-enhanced algorithm development 89%
- Evaluating algorithmic fairness in the presence of clinical guidelines: the case of atherosclerotic cardiovascular disease risk estimation 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.