Back

Predicting Hypertension Among HIV Patients on Antiretroviral Therapy in Rural Eastern Cape, South Africa Using Machine Learning

Tsuro, U.; Ncube, T.; Oladimeji, K. E.; Apalata, T. R.

2025-01-12 health informatics
10.1101/2025.01.11.25320377 medRxiv
Show abstract

BackgroundHypertension continues to be a major challenge in developing countries like South Africa, as it significantly contributes to the cardiovascular disease burden in these countries. This study aimed to utilize the machine learning (ML) models to anticipate the incidence of hypertension in HIV patients under antiretroviral therapy (ART) in rural Eastern Cape, South Africa. MethodsThis research carried out a retrospective cohort study and created and tested six machine learning algorithms: Neural Networks, Random Forest, Logistic Regression, Naive Bayes, K-Nearest Neighbours and XGBoost. The goal was to predict the likelihood of developing hypertension. Feature selection was done using the Boruta method and the model was assessed using several metrics including aiming, precision, recall, F1 score, and area under the receiver operating characteristic curve (AUC). ResultsXGBoost outperformed all other models with an AUC of 0.96, which further suggests it can effectively distinguish between hypertensives and normotensives. In the case of Boruta analysis, some aggravated risk factors were age category, time on ART, BMI category, waist to hip ratio, waist size, family history of HBP and relationship status, physical activity, LDL cholesterol level, awareness of high blood pressure, education level, use of ART and diabetes mellitus. ConclusionsThis study has highlighted the utility of XGBoost, as one of the advanced machine learning algorithms, in reliably forecasting the occurrence of hypertension in HIV ART patients in a rural setting. The established risk factors elucidate the complexity behind the hypertension emergence and hence the need for triad approaches which include lifestyle changes, clinical treatments, and demographic solutions to tackle the public health problem.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.