Large datasets from Electronic Health Records predict seizures after ischemic strokes: A Machine Learning approach.
Lekoubou, A.; Petucci, j.; Ajala, T. F.; Katoch, A.; Sen, S.; Honavar, V.
Show abstract
ObjectiveTo develop an artificial intelligence, machine learning prediction model for estimating the risk of seizures 1 year and 5 years after ischemic stroke (IS) using a large dataset from Electronic Health Records. BackgroundSeizures are frequent after ischemic strokes and are associated with increased mortality, poor functional outcomes, and lower quality of life. Separating patients at high risk of seizures from those at low risk of seizures is needed for treatment and clinical trial planning, but remains challenging. Machine learning (ML) is a potential approach to solve this paradigm. Design/MethodsWe identified patients (aged [≥]18 years) with IS without a prior diagnosis of seizures from 2015 until inception (08/09/22) in the TriNetX Research Network, using the International Classification of Diseases, Tenth Revision (ICD-10) I63, excluding I63.6 (venous infarction). The outcome of interest was any ICD-10 diagnosis of seizures (G40/G41) at 1 year and 5 years following the index IS. We applied a conventional logistic regression and a Light Gradient Boosted Machine algorithm to predict the risk of seizures at 1 year and 5 years. The performance of the model was assessed using the area under the receiver operating characteristics (AUROC), the area under the precision-recall curve (AUPRC), F1 statistic, model accuracy, balanced accuracy, precision, and recall, with and without anti-seizure medication use in the models. ResultsOur study cohort included 430,254 IS patients. Seizures were present in 18,502 (4.3%) and (5.3%) patients within 1 and 5 years after IS, respectively. At 1-year, the AUROC, AUPRC, F1 statistic, accuracy, balanced-accuracy, precision, and recall were respectively 0.7854 (standard error: 0.0038), 0.2426 (0.0048), 0.2299 (0.0034), 0.8236 (0.001), 0.7226 (0.0049), 0.1415 (0.0021), and 0.6122, (0.0095). Corresponding metrics at 5 years were 0.7607 (0.0031), 0.247 (0.0064), 0.2441 (0.0032), 0.8125 (0.0013), 0.7001 (0.0045), 0.155 (0.002) and 0.5745 (0.0095). ConclusionOur findings suggest that ML models show good model performance for predicting seizures after IS.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Short-term Functional Outcomes of Patients with acute intracerebral hemorrhage in the Native and Expatriate Population 94%
- Automated Identification of Thrombectomy Amenable Vessel Occlusion on Computed Tomography Angiography using Deep Learning 93%
- Left Hemisphere Bias of NIH Stroke Scale is Most Severe for Middle Cerebral Artery Strokes 93%
Similar papers in this journal
- Modified Rankin Scale Disability Status at Day 4 Poststroke is an Informative Predictor of Long-Term Day 90 Outcome 95%
- Time is Brain: Detection of Nonconvulsive Seizures and Status Epilepticus During Acute Stroke Evaluation Using Point-of-Care Electroencephalography 95%
- White matter lesions as a prognostic marker of recurrence in cryptogenic stroke with high-risk patent foramen ovale 93%
Similar papers in this journal
- Intracerebral Hemorrhage Outcomes after Reversal of Subtherapeutic Warfarin: Analysis of Data from GWTG-Stroke 94%
- Prospective Observational Cohort Study Of Tenecteplase Versus Alteplase In Routine Clinical Practice 94%
- Effect of Time to Thrombolysis on Clinical Outcomes in Patients with Acute Ischemic Stroke Treated with Tenecteplase Compared to Alteplase: Analysis from the AcT Randomized Controlled Trial 94%
Similar papers in this journal
Similar papers in this journal
- Leveraging Machine Learning for Enhanced and Interpretable Risk Prediction of Venous Thromboembolism in Acute Ischemic Stroke Care 96%
- Non-invasive Auricular Vagus nerve stimulation for Subarachnoid Hemorrhage (NAVSaH): Protocol for a prospective, triple-blinded, randomized controlled trial 93%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.