Back

Re: Machine Learning Approaches to Identify Communities with High HIV Prevalence in Resource-Limited Settings using Social, Economic and Behavioral Data

Milali, M. P.; Assefa, F. B.; Gathungu, D. K.; Mwalili, S.; Nyimbili, S.; Sivile, S.; Mulenga, L.; Braithwaithe, R. S.; Cuadros, D. F.; Bershteyn, A.

2025-11-14 hiv aids
10.1101/2025.11.10.25339949 medRxiv
Show abstract

BackgroundIdentifying communities with high HIV prevalence is crucial for public health officials, researchers, and policymakers to effectively monitor the epidemic and evaluate interventions. Population-based HIV biomarker surveys face logistical challenges such as cost, need for personnel trained in specimen collection, specimen transport and processing, and participant reluctance to test due to factors such as stigma, history of recent testing, and the perception of being at low risk for HIV infection. This study explores the potential of identifying communities with high HIV prevalence using socio-economic, behavioral, and other community-level data in the absence of direct HIV biomarkers. MethodUsing the methods of Partial Least Squares (PLS) and Random Forests (RF), we developed machine learning models to predict HIV prevalence based on socio-economic and behavioral variables from Population-based HIV Impact Assessments (PHIA) surveys. Community HIV prevalence, derived from the PHIA biomarkers dataset, served as the dependent variable. Initially, models were trained to classify communities into <10% or [&ge;]10% HIV prevalence categories. This procedure was repeated for prevalence thresholds of 5%, 7%, 15%, and 20%. ResultsPLS and RF achieved 79% and 80.5% accuracy, respectively, in classifying communities as having higher or lower HIV prevalence at a 10% threshold. At the 5%, 7%, and 15% thresholds, the models achieved similar accuracies, demonstrating consistent performance across varying thresholds, with RF slightly outperforming PLS. In both models, the variables that contributed most to classification included having a first intercourse experience before age 15, being uncircumcised, having a history of not using condoms, being in the lowest wealth quintile, experiencing physical or sexual violence, and having extramarital partners. ConclusionsThe study demonstrates that socioeconomic and behavioral variables can effectively predict community-level HIV prevalence using machine learning models. These insights have the potential to guide the distribution of HIV resources, particularly where direct community testing is infeasible, and to enhance understanding of the HIV epidemics.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.