Back

Heart Failure Prediction: An Explainable Cross-Validated Comparison of Several Machine Learning Models.

ALI, M. S.; Islam, T.

2025-10-22 health informatics
10.1101/2025.10.21.25338436 medRxiv
Show abstract

Heart disease is the leading cause of death worldwide, contributing to millions of fatalities each year. Early detection and accurate risk prediction are therefore critical for timely intervention and improved patient outcomes. In this study, seven different Machine Learning (ML) prediction models were tested on the collections of heart disease datasets. The data was divided, with 80% used to train the models and 20% saved to test them later. The settings for the models were carefully adjusted using a method that checks their performance multiple times to find the best ones. The models were then evaluated on the unseen test data. Their performance was measured using several metrics (Accuracy, Recall, Precision, PR-AUC and ROC-AUC), and their results were further examined using a confusion matrix. In addition to traditional evaluation metrics, SHapley Additive exPlanations (SHAP) analysis was employed to interpret the contribution of each feature to the models predictions. It was observed that the Multi-Layer Perceptron (MLP) achieved the highest performance on both datasets, demonstrating strong predictive capability while remaining interpretable through the integration of SHAP. This study shows that modern ML models can be very good at predicting heart disease risk, and provide explainable performance. This study provides an effective approach for predicting heart disease risk with an explainable model to help doctors choose the best tool.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.