Machine Learning Models for Predicting Stroke Risk Among Patients with Coronary Heart Disease
Wanyonyi, M.; Kitavi, D. M.; Musyoka, F. M.; Morris, Z. N.
Show abstract
Stroke is a major global health burden and a frequent severe complication among patients with coronary heart disease. Early identification of individuals at high risk is essential for prevention; however, conventional clinical models often fail to capture the complex interactions underlying stroke risk in this population. This study developed an integrated machine learning framework to predict stroke risk using a large real-world dataset. Multiple algorithms were evaluated, including logistic regression, decision trees, random forests, support vector machines, naive Bayes, multilayer perceptrons, LightGBM, XGBoost, deep learning models, and a stacked ensemble. Class imbalance was addressed using stratified sampling and synthetic minority oversampling. The stacked ensemble demonstrated superior performance, achieving an AUC-ROC of 0.96, precision of 0.91, recall of 0.89, and a Matthews correlation coefficient of 0.82. LightGBM and XGBoost also performed strongly, with AUC-ROC values of 0.94 and 0.95, respectively, and low inference latency. To enhance clinical interpretability, explainable AI techniques (SHAP and LIME) were applied, identifying key risk factors such as prior heart attack, body mass index, age, and lifestyle behaviors. These findings indicate that machine learning models can substantially improve stroke risk prediction in patients with coronary heart disease, supporting clinically actionable and scalable decision-support systems.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhanced machine learning and hybrid ensemble approaches for coronary heart disease prediction 96%
- Leveraging Machine Learning for Enhanced and Interpretable Risk Prediction of Venous Thromboembolism in Acute Ischemic Stroke Care 95%
- A machine learning approach to identifying important features for achieving step thresholds in individuals with chronic stroke 95%
Similar papers in this journal
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 98%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 95%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 95%
Similar papers in this journal
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 95%
- Electrocardiogram Analysis of Post-Stroke Elderly People Using One-dimensional Convolutional Neural Network Model with Gradient-weighted Class Activation Mapping 94%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 94%
Similar papers in this journal
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 95%
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 95%
- Improving irregular temporal modeling by integrating synthetic data to the electronic medical record using conditional GANs: a case study of fluid overload prediction in the intensive care unit 94%
Similar papers in this journal
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 95%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.