Benchmarking Transformer-Based and Conventional Machine Learning Models for Cardiovascular Disease Prediction on Datasets of Varying Scale and Complexity
Upadhyayula, S. K.; Pothugunta, R.
Show abstract
ObjectiveCardiovascular Diseases (CVDs) remain the leading cause of death worldwide, creating an urgent need for accurate risk prediction. Machine learning (ML) methods are well established, but transformer-based deep learning architectures are emerging as promising alternatives. Their comparative value, particularly under challenges such as class imbalance, is still unclear. MethodsWe systematically compared transformer models (FT-Transformer, SAINT, TabNet, TabTransformer) with conventional ML algorithms (support vector machine, random forest, XGBoost, etc) using three public CVD datasets of increasing size and complexity: the balanced UCI dataset, the imbalanced Framingham dataset, and the large-scale Kaggle dataset. A consistent preprocessing pipeline was applied, with MICE imputation for missing data and SMOTETomek resampling for imbalance for the Framingham dataset. Models were assessed with stratified 10-fold cross-validation, and their performance was statistically compared across datasets. Explainability was explored using SHAP feature importance. ResultsPerformance varied with dataset characteristics. On the small, balanced UCI dataset, FT-Transformer achieved near-perfect accuracy (AUC > 0.99), comparable to XGBoost and random forest. On the imbalanced Framingham dataset, sensitivity remained low overall, though FT-Transformer achieved the best trade-off. On the Kaggle dataset, FT-Transformer and XGBoost performed similarly, both identifying systolic blood pressure and age as major predictors. ConclusionTransformer models show strong potential for structured health data but remain sensitive to imbalance, where conventional ML retains advantages. Careful dataset-aware model selection is essential for CVD prediction.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 95%
- Multiple Instance Learning Framework can Facilitate Explainability in Murmur Detection 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
Similar papers in this journal
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 95%
- Advancing Cardiovascular Disease Diagnosis: A Robust ML Ecosystem Integrating Early Detection, Responsible AI Framework, and Causal Inference 95%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 95%
Similar papers in this journal
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 95%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 95%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 94%
Similar papers in this journal
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 96%
- Deep Neural Survival Networks for Cardiovascular Risk Prediction: The Multi-Ethnic Study of Atherosclerosis (MESA) 94%
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.