Limited Predictability of Client Attendance in a Support Program for HIV Vertical Transmission Prevention: A Comparison of Machine Learning and Community Health Worker Predictions
Olckers, M.; Lam, A.; Makhupula, L.; Mvubu, M.
Show abstract
Client attendance is vital for the success of HIV vertical transmission prevention programs, yet 23.4% of clients missed follow-up appointments after enrolling in a community health worker-led program (n=24,807, Aug-Dec 2022). Predicting which clients are most likely to miss appointments could enable targeted interventions to improve retention. While machine learning appears well-suited for this prediction task, its effectiveness compared to community health worker judgment remains unexplored. We evaluated three machine learning approaches--logistic regression, balanced random forest, and gradient-boosted trees--trained on client enrollment records (n=51,297 training; n=18,577 test) and compared their performance to predictions by community health workers (n=61), who possess direct client interactions and contextual insights. Machine learning models achieved modest predictive performance, with the balanced random forest showing the best accuracy (ROC AUC=0.689). Community health worker predictions similarly showed low accuracy despite their rich contextual knowledge, suggesting attendance is fundamentally difficult to predict. Qualitative insights identified complex and dynamic barriers including stigma, transportation difficulties, and competing personal commitments, often unpredictable at enrollment. These findings highlight fundamental limitations in predicting client attendance and suggest that merely accumulating additional data might not enhance predictive accuracy. Instead, resources might be better allocated to addressing systemic barriers identified by community health workers.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Impact of a pilot mHealth intervention on treatment outcomes of TB patients seeking care in the private sector using Propensity Scores Matching – Evidence collated from New Delhi, India 96%
- Combining participatory mapping and route optimization algorithms to inform the delivery of community health interventions at the last mile 94%
- Developing contents for a digital drug adherence tool with reminder cues and personalized feedback: a formative mixed-methods study among children and adolescents living with HIV in Tanzania 92%
Similar papers in this journal
- Health worker acceptability of an HIV testing mobile health application within a rural Zambian HIV treatment programme 94%
- Cost savings in male circumcision post-operative care continuum in rural and urban South Africa: Evidence on the importance of initial counselling and daily SMS 94%
- “It reminds me and motivates me” : Human-centered design and implementation of an interactive, SMS-based digital intervention to improve early retention on antiretroviral therapy: usability and acceptability among new initiates in a high-volume, public clinic in Malawi 93%
Similar papers in this journal
Similar papers in this journal
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 91%
- Machine Learning Directed Interventions Associate with Decreased Hospitalization Rates in Hemodialysis Patients 90%
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 90%
Similar papers in this journal
- Collecting mortality data via mobile phone surveys: a non-inferiority randomized trial in Malawi 94%
- Improving patient-centred counselling skills among lay healthcare workers in South Africa using the Thusa-Thuso motivational interviewing training and support program 94%
- Examining National Health Insurance Fund Members’ preferences and trade-offs for the attributes of contracted outpatient facilities in Kenya: a discrete choice experiment 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.