Developing Predictive Algorithms for Patient Retention Using Machine Learning and Deep Learning to Improve HIV Care in Uganda
Mirugwe, A.; Ssevvume, S.; Fitzmaurice, A. G.; Mpango, J.; Namale, A.; Akello, E.; Katongole, P.; Muhumuza, S.; Mayanja, N.; Mbaka, P.; Sande, E.; Musenge, K.
Show abstract
BackgroundAchieving high retention of people living with HIV (PLHIV) in care remains a challenge in Uganda, despite substantial progress towards UNAIDS 95-95-95 targets. This study used advanced machine learning and deep learning techniques applied to de-identified longitudinal PLHIV data routinely collected in HIV clinics in Uganda to predict clients at high risk of missing treatment appointments. MethodsWe compared the performance of traditional machine learning models (i.e., Decision Tree, Random Forest, AdaBoost, and XGBoost) and the Bidirectional Encoder Representations from Transformers (BERT) model, which is more suitable for analyzing longitudinal data. Feature importance using the Shapley additive exPlanations method was used to identify the most influential predictors. We also evaluated the impact of various sampling techniques, i.e., undersampling, oversampling, and synthetic minority oversampling, to address class imbalance and improve model performance. Model performance was evaluated using accuracy, precision, recall, F1-score, and the Area Under the Curve-Receiver Operating Characteristic (AUC-ROC) metrics. ResultsThe study was based on a longitudinal dataset of 66,206 PLHIV who initiated HIV care during 2000-2023 in 86 health facilities. The data comprised 1,479,121 clinical visits, an average of 22 clinical visits; 158,266 (10.7%) missed appointments, and 49,588 (74.9%) of clients missed at least one appointment. Median (interquartile range [IQR]) age was 36.0 [29.0 - 47.0] years, and the majority (n=43132, 65%) were female. The BERT model demonstrated superior performance, achieving an AUC score of 0.96, 94.8% accuracy, 97.1% precision, 100% recall, and an F1-score of 94.2%. In comparison, the XGBoost model with undersampling achieved an AUC score of 0.90, 80.7% accuracy, 97.1% precision, 80.8% recall, and an F1-score of 88.2%. Feature importance analysis showed that treatment adherence, visit frequency, treatment duration, and visits on the current regimen were the most influential predictors of appointment interruption. ConclusionThis study highlights the efficacy of transformer-based models like BERT in handling longitudinal clinical data and improving patient retention predictions. Integrating these predictive models into electronic medical systems will facilitate proactive treatment strategies, enabling the identification of clients at risk of disengagement before they miss appointments. This approach may contribute to the improvement of HIV care and support progress towards achieving the HIV program targets in Uganda and potentially elsewhere.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Impact of a pilot mHealth intervention on treatment outcomes of TB patients seeking care in the private sector using Propensity Scores Matching – Evidence collated from New Delhi, India 95%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 94%
- External validation of a paediatric SMART triage model for use in resource limited facilities 92%
Similar papers in this journal
- A Machine Learning-Based Prediction of Hospital Mortality in Mechanically Ventilated ICU Patients 94%
- Identification of high-risk COVID-19 patients using machine learning 93%
- Isoniazid preventive therapy completion and factors associated with non-completion among patients on antiretroviral therapy at Kisenyi Health Centre IV, Kampala, Uganda 93%
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 93%
- Predicting mortality in SARS-COV-2 (COVID-19) positive patients in the inpatient setting using a Novel Deep Neural Network 92%
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.