Real-World Validation of Machine Learning Models for HIV Treatment Adherence Prediction and Care Gap Quantification: A Multi-Country Analysis of 192,732 Clinical Records
Chinthala, L. K.
Show abstract
Delayed diagnosis and poor antiretroviral therapy (ART) adherence remain primary drivers of HIV-related morbidity in low-resource settings, yet real-world AI validation at scale is lacking. We conducted a retrospective validation study using two publicly available, de-identified datasets: a Quality of Care cohort of 27,288 HIV-positive patients on ART across multiple healthcare facilities, and the CEPHIA multi-country assay database comprising 165,444 specimen records from six countries. Four machine learning classifiers were evaluated using 10-fold stratified cross-validation with SMOTE applied strictly to training folds. Explicit data leakage prevention, ablation analysis, calibration assessment, and bootstrap confidence intervals were applied. Economic projections used one-way sensitivity analysis. This study adheres to TRIPOD reporting guidelines. Random Forest achieved AUC-ROC of 0.9753 (95% CI: 0.970-0.975), sensitivity 87.3% (95% CI: 86.4-88.2%), specificity 95.7% (95% CI: 95.2-96.2%), and Brier score 0.079. Ablation testing confirmed robustness (AUC 0.963 without the primary predictor). Temporal validation on held-out future patients yielded AUC 0.772 (95% CI: 0.744-0.802), confirming generalisation across time. Real-world analysis revealed median diagnosis-to-ART delay of 74 days, with 47.3% of patients exceeding 90 days and 36.7% presenting with CD4 below 200 cells per microlitre. Multi-country CEPHIA analysis identified 18.6% HIV recency within the 130-day early-intervention window. Decision curve analysis confirmed net clinical benefit across threshold probabilities 0.03-0.45. Subgroup analysis demonstrated consistent AUC across sex, age, CD4 strata, and WHO staging (max difference 0.051). Economic modelling projected base-case savings of USD 415 per patient (USD 2.07 million per 5,000-patient cohort). These findings provide large-scale empirical evidence that AI-driven informatics can predict ART adherence failure and quantify systemic care gaps, offering a scalable framework for equitable HIV care delivery in resource-limited settings. Prospective external validation is required before clinical deployment.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Digitising point of care HIV test results to accurately measure, and improve performance towards, the UNAIDS 90-90-90 targets 94%
- The Risks and Benefits of Providing HIV Services during the COVID-19 Pandemic 94%
- Impact Of The COVID-19 Pandemic On Routine HIV Care And Antiretroviral Treatment Outcomes In Kenya: A Nationally Representative Analysis 94%
Similar papers in this journal
- National HIV testing and diagnosis coverage in sub-Saharan Africa: a new modeling tool for estimating the \"first 90\" from program and survey data 94%
- Use of undetectable viral load to improve survey estimates of known HIV-positive status and antiretroviral treatment coverage 93%
- A Decision Analytics Model to Optimize Investment in Interventions Targeting the PrEP Cascade of Care 93%
Similar papers in this journal
- Field performance and cost-effectiveness of a point-of-care triage test for HIV virological failure in Southern Africa 96%
- Estimating the impact of disruptions due to COVID-19 on HIV transmission and control among men who have sex with men in China 93%
- Diagnostic accuracy of a point-of-care urine tenofovir assay, and associations with HIV viraemia and drug resistance among people receiving dolutegravir and efavirenz-based antiretroviral therapy 93%
Similar papers in this journal
- Machine learning models predict long COVID outcomes based on baseline clinical and immunologic factors 90%
- Predicting future hospital antimicrobial resistance prevalence using machine learning 89%
- A significant increase in Tuberculosis diagnosis is required to mitigate the impact of COVID-19 on its future burden 89%
Similar papers in this journal
- Use of HIV Recency Assays for HIV Incidence Estimation and Non-Incidence Surveillance Use Cases: A systematic review 93%
- Uncovering clinical risk factors and prediction of severe COVID-19: A machine learning approach based on UK Biobank data 90%
- Toward Using Twitter for PrEP-Related Interventions: An Automated Natural Language Processing Pipeline for Identifying Gay or Bisexual Men in the United States 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.