Comparative Analysis of Machine Learning Models vs. Traditional Clinical Calculators for Cardiovascular Risk Prediction
Arango Plaza, N.; Salcedo Echeverry, G. E.; Arenas-Soto, A. F.
Show abstract
Background: Cardiovascular diseases (CVD) remain the leading global cause of mortality, responsible for approximately 31% of all deaths worldwide in 2021. Traditional risk calculators, including Framingham, ASCVD, SCORE, and SCORE2, have long constituted the cornerstone of primary prevention strategies; however, they were derived predominantly from high-income European and North American populations, thereby limiting their predictive accuracy in diverse epidemiological contexts, particularly among Hispanic/Latino communities. Machine learning (ML) offers an alternative to capture the non-linear interactions inherent in biomedical data. Objective: The present study develops and validates ML-based models for cardiovascular mortality prediction using the National Health and Nutrition Examination Survey (NHANES) 1999-2018 dataset, and systematically compares their discriminative performance against eleven conventional clinical CVD risk calculators. Materials and Methods: A dedicated software platform, "CardioPrediQ," was designed to integrate multiple CVD calculators with ML-based risk assessment. A cohort of 12,847 participants with 16 predictor variables was derived from NHANES. Six algorithms (Logistic Regression, Cox Proportional Hazards, Gradient Boosting, AdaBoost, Random Forest, and Extra Trees) were trained in combination with six class-balancing strategies, yielding 36 model configurations. All models were trained on a stratified 70/30 split and calibrated using the Saerens prior probability adjustment method. Performance was evaluated using AUC-ROC, sensitivity, specificity, F1-score, and a weighted composite score. DeLong's test was employed to assess the statistical significance of AUC differences between the best-performing ML model and each conventional calculator. Results: Gradient Boosting with 2:1 oversampling and Saerens calibration achieved the best overall performance (AUC = 0.8934; composite score = 0.7904), outperforming all traditional calculators in composite ranking. The top six positions were occupied exclusively by ML and statistical models. The mean age of cardiovascular decedents was 67.43 years compared with 47.74 years among survivors. DeLong's test confirmed statistical superiority over six traditional CVD calculators (p < 0.05), whereas the difference against the top-performing calculators (ASCVD, HEARTS Caribbean, ASCVD Colombia, SCORE2, HEARTS North America) did not reach statistical significance. Age dominated feature importance at 41.2% relative weight, followed by systolic blood pressure (18.7%). Saerens calibration reduced the Brier score from 0.1286 to 0.1158, substantially improving probability calibration. Conclusions: ML models demonstrated superior composite performance over traditional calculators. The statistical equivalence with the highest-performing conventional calculators in the NHANES cohort is context-dependent and validates the methodological pipeline. The CardioPrediQ platform addresses the critical need for integrated, scalable CVD risk assessment tools, which is particularly relevant for Latin American populations where calculator validation remains limited. These findings support the integration of calibrated ML-based risk prediction into clinical practice while underscoring the importance of probability calibration for informed clinical decision-making.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and validation of a cardiovascular diseases risk prediction model for Chinese males (CVDMCM) 95%
- Predicting long-term prognosis after percutaneous coronary intervention in patients with acute coronary syndromes: a prospective nested case-control analysis for county-level health services 95%
- Fairness in Cardiac Magnetic Resonance Imaging: Assessing sex and racial bias in deep learning-based segmentation 93%
Similar papers in this journal
- Actionable absolute risk prediction of atherosclerotic cardiovascular disease: a behavior-management approach based on data from 464,547 UK Biobank participants 96%
- Enhanced machine learning and hybrid ensemble approaches for coronary heart disease prediction 94%
- ChatGPT Provides Inconsistent Risk-Stratification of Patients With Atraumatic Chest Pain 94%
Similar papers in this journal
- Using ECG Machine Learning for Detection of Cardiovascular Disease in African American Men and Women: the Jackson Heart Study 96%
- International Evaluation Of An Artificial Intelligence-Powered Ecg Model Detecting Occlusion Myocardial Infarction 94%
- Prediction of Cardiovascular Markers and Diseases Using Retinal Fundus Images and Deep Learning: A Systematic Scoping Review 93%
Similar papers in this journal
- Walking pace optimizes conventional cardiovascular disease risk prediction models among vulnerable subpopulations: a prospective cohort study 94%
- Artificial intelligence of arterial Doppler waveforms to predict major adverse outcomes among patients evaluated for peripheral artery disease 94%
- Age-stratified Prevalence and Relative Prognostic Significance of Traditional Atherosclerotic Risk Factors: A Report from the Nationwide Registry of Percutaneous Coronary Interventions in Japan 94%
Similar papers in this journal
- Development and validation of a risk prediction algorithm for high-risk populations combining genetic and conventional risk factors of cardiovascular disease 95%
- Lipoprotein(a) and cardiovascular disease: prediction, attributable risk fraction and estimating benefits from novel interventions 93%
- Artificial Intelligence Methods to Detect Heart Failure with Preserved Ejection Fraction (AIM-HFpEF) within Electronic Health Records: An equitable disease prediction model 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.