Back

Benchmarking Transformer-Based and Conventional Machine Learning Models for Cardiovascular Disease Prediction on Datasets of Varying Scale and Complexity

Upadhyayula, S. K.; Pothugunta, R.

2025-08-11 health informatics
10.1101/2025.08.03.25332878 medRxiv
Show abstract

ObjectiveCardiovascular Diseases (CVDs) remain the leading cause of death worldwide, creating an urgent need for accurate risk prediction. Machine learning (ML) methods are well established, but transformer-based deep learning architectures are emerging as promising alternatives. Their comparative value, particularly under challenges such as class imbalance, is still unclear. MethodsWe systematically compared transformer models (FT-Transformer, SAINT, TabNet, TabTransformer) with conventional ML algorithms (support vector machine, random forest, XGBoost, etc) using three public CVD datasets of increasing size and complexity: the balanced UCI dataset, the imbalanced Framingham dataset, and the large-scale Kaggle dataset. A consistent preprocessing pipeline was applied, with MICE imputation for missing data and SMOTETomek resampling for imbalance for the Framingham dataset. Models were assessed with stratified 10-fold cross-validation, and their performance was statistically compared across datasets. Explainability was explored using SHAP feature importance. ResultsPerformance varied with dataset characteristics. On the small, balanced UCI dataset, FT-Transformer achieved near-perfect accuracy (AUC > 0.99), comparable to XGBoost and random forest. On the imbalanced Framingham dataset, sensitivity remained low overall, though FT-Transformer achieved the best trade-off. On the Kaggle dataset, FT-Transformer and XGBoost performed similarly, both identifying systolic blood pressure and age as major predictors. ConclusionTransformer models show strong potential for structured health data but remain sensitive to imbalance, where conventional ML retains advantages. Careful dataset-aware model selection is essential for CVD prediction.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.