Predicting Acute Cerebrovascular Events in Stroke Alerts Using Large-Language Models and Structured Data
Erekat, A.; Downes, M. H.; Stein, L. K.; Delman, B. N.; Karp, A. M.; Tripathi, A.; Nadkarni, G. N.; Kupersmith, M. J.; Kummer, B. R.
Show abstract
BackgroundAcute stroke alerts are often activated for non-cerebrovascular conditions, leading to false positives that strain clinical resources and promote diagnostic uncertainty. We sought to develop machine learning (ML) models integrating large-language models (LLMs), structured electronic health record data, and clinical time series data to predict the presence of acute cerebrovascular disease (ACD) at stroke alert activation. MethodsWe derived a series of ML models using retrospective data from stroke alerts activated at Mount Sinai Health System between 2011 and 2021. We extracted structured data (demographics, medical comorbidities, medications, and engineered time-series features from vital signs and lab results) as well as unstructured clinical notes available prior to the time of stroke alert. We processed clinical notes using three embedding approaches: word embeddngs, biomedical embeddings (BioWordVec), and LLMs. Using a radiographic gold standard for acute intracranial vascular event, we used an auto-ML approach to train one model based on unstructured data and five models based on different combinations of structured data. We evaluated models individually using the area under the receiver operating characteristic curve (AUROC), mean positive predictive value (PPV), sensitivity, and F1-score. We then combined the 6 model logits into a multimodal ensemble by weighting their logits based on F1-score, determining ensemble performance using the same metrics. ResultsWe identified 16,512 stroke alerts corresponding to 14,233 unique patients over the study period, of which 9,013 (54.6%) were due to ACD. The multi-modal model (AUROC 0.72, PPV 0.68, sensitivity 0.76, F1 0.72) outperformed all individual models by AUROC. One structured model based on demographics, comorbidities, and medications demonstrated the highest sensitivity (0.95). ConclusionsWe developed a multi-modal ML model to predict ACD at stroke alert activation. This approach has promise to optimize stroke triage and reduce false-positive activations.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 93%
- Machine learning-based forecasting of daily acute ischemic stroke admissions using weather data 93%
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 94%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 92%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 92%
Similar papers in this journal
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 94%
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 94%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.