Development of a Hypertension Risk Prediction Model using Nationally Representative Survey Data: A Machine Learning Approach and Web Application Deployment
Pandey, S.; Pandey, A.; Neupane, A.; Subedi, D. J.; Guragain, A.
Show abstract
BackgroundHypertension is a major modifiable risk factor for cardiovascular diseases. Early identification of high-risk individuals using predictive models can facilitate targeted interventions. This study aims to develop and validate a machine learning-based risk prediction model for hypertension using complex, nationally representative survey data and to deploy it as an accessible web application. MethodsWe utilized data from the WHO Steps Survey Nepal, 2019 including 5,593 participants. The dataset featured clustering, stratification, and sampling weights, which were incorporated into the analysis. Fourteen initial predictors were considered. We employed a combination of SMOTENC and KMeansSMOTE to address class imbalance and optimized prediction thresholds for each model. Six machine learning algorithms (Logistic Regression, Naive Bayes, Random Forest, LightGBM, XGBoost, and SVM) were trained and evaluated based on AUC ROC, precision-recall, calibration, and clinical utility (Decision Curve Analysis). Model interpretability was assessed using SHAP values. The best model was deployed as an interactive web application using Streamlit. ResultsThe prevalence of hypertension in the weighted sample was 26.1%. After feature selection using SHAP analysis, seven key predictors were retained: age, smoking, waist-hip ratio risk, heavy alcohol use, physical activity, fasting blood sugar, and total cholesterol. Logistic Regression demonstrated the best overall performance (AUC ROC: 0.718, F1-Score: 0.552) and was well-calibrated. It also offered the highest net benefit across a wide range of clinical thresholds. The model was successfully deployed as a publicly available web application (https://htnrisknepal.streamlit.app/). ConclusionsWe developed a robust, interpretable, and clinically useful hypertension risk prediction model. The deployment of the model as an open-access web application bridges the gap between research and practical implementation, enabling its use by healthcare workers and for public health screening initiatives. Our methodology provides a reliable framework for building and deploying predictive models from public health survey data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automated Image Transcription for Perinatal Blood Pressure Monitoring Using Mobile Health Technology 94%
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 93%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 92%
Similar papers in this journal
- Empowering Healthcare Professionals in West Africa □ A Feasibility Study and Qualitative Assessment of a Dietary Screening Tool to Identify Adults at High Risk of Hypertension 94%
- Comparison of WHO laboratory-based and non-laboratory-based CVD risk Charts among Hypertensive Adults Attending Primary Healthcare Centers in West Africa Sub-region 94%
- bp: Blood Pressure Analysis in R 94%
Similar papers in this journal
- Social support and ideal cardiovascular health in urban Jamaica: a cross-sectional study 92%
- Access to hypertension services and health-seeking experiences in rural Coastal Kenya: A qualitative study 92%
- Regional Prevalence of Hypertension Among People Diagnosed with Diabetes in Africa, A Systematic Review and Meta-analysis 92%
Similar papers in this journal
- Causal modeling of chronic kidney disease in a participatory framework for informing the inclusion of social drivers in health algorithms 92%
- Learning Decision Thresholds for Risk-Stratification Models from Aggregate Clinician Behavior 92%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 92%
Similar papers in this journal
- Can machine learning improve risk prediction of incident hypertension? An internal method comparison and external validation of the Framingham risk model using HUNT Study data 96%
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 94%
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.