Comprehensive, Transparent, and Fair Machine Learning Models for Hypertension Risk Prediction: Benchmarking With Framingham, External Validation, Individual-Level Analysis, and Equitable Clinical Utility
Amirian, P.; Zarpoosh, M.
Show abstract
BackgroundHypertension (HTN) is a leading, yet often underdiagnosed, cause of cardiovascular diseases worldwide. While clinical risk scores like the Framingham Risk Score (FRS) are commonly used, their limited capacity to capture complex risk patterns and their poor generalizability hinder optimal prediction and prevention in resource-limited settings. MethodsWe developed and rigorously validated a comprehensive suite of machine learning (ML) models, including tree-based, linear, kernel, neural network, and ensemble classifiers for incident HTN prediction using data from our internal cohort (n=8,054, Iran). All models were benchmarked against the FRS. External validation was conducted on the NHANES cohort (n=6,266, USA), employing harmonized features. We systematically assessed model performance across multiple metrics (ROC AUC, PR AUC, F1, and Brier), evaluated calibration, clinical decision benefit, and learning curves. We conducted extensive subgroup and fairness analyses by sex, education, and socioeconomic status. Interpretability was ensured via SHAP values and permutation importance. Additionally, individualized counterfactual analyses, bootstrap prediction intervals, and patient-level risk vignettes were provided. ResultsML models, especially our ensembles and LightGBM, significantly outperformed FRS in discrimination (mean ROC AUC improvement up to 0.04, 95% CI not crossing zero), with robust generalizability confirmed in external validation. All models demonstrated minimal performance disparities across demographic and socioeconomic subgroups, and SHAP analyses identified SBP, age, and BMI as consistently strong predictors. Individualized risk vignettes illustrated the clinical nuance of ML predictions compared to categorical clinical scores. ConclusionThis open, reproducible study demonstrates that modern, interpretable ML models provide significant and equitable improvements over standard clinical risk scores for HTN prediction, with strong external validity and clinical utility. An online risk calculator is provided to facilitate real-world deployment.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 95%
- Improving Pre-eclampsia Risk Prediction by Modeling Individualized Pregnancy Trajectories Derived from Routinely Collected Electronic Medical Record Data 94%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 93%
Similar papers in this journal
- Actionable absolute risk prediction of atherosclerotic cardiovascular disease: a behavior-management approach based on data from 464,547 UK Biobank participants 96%
- Development and Validation of the Michigan Chronic Disease Simulation Model (MICROSIM) 94%
- An assessment of the value of deep neural networks in genetic risk prediction for surgically relevant outcomes 93%
Similar papers in this journal
- Sociodemographic Characteristics and Longitudinal Progression of Multimorbidity: A Multistate Modelling Analysis of a Large Primary Care Records Dataset in England 92%
- Genetically-proxied therapeutic inhibition of antihypertensive drug targets and risk of common cancers 91%
- Cost-effectiveness of leveraging existing HIV primary health systems and community health workers for hypertension screening and treatment in Africa: an individual-based modelling study 91%
Similar papers in this journal
- Can machine learning improve risk prediction of incident hypertension? An internal method comparison and external validation of the Framingham risk model using HUNT Study data 97%
- Nationwide prediction of type 2 diabetes comorbidities 95%
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 94%
Similar papers in this journal
- Contrasting factors associated with COVID-19-related ICU admission and death outcomes in hospitalised patients by means of Shapley values 93%
- Identification of a serum proteomic biomarker panel using diagnosis specific ensemble learning and symptoms for early pancreatic cancer detection 91%
- Explainable deep transfer learning model for disease risk prediction using high-dimensional genomic data 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.