Back

A Rule-Based Machine Learning Model for Predicting Virological Failure Among Children Living With HIV in Malawi

Chiphe, C.

2026-03-10 hiv aids
10.64898/2026.03.09.26347945 medRxiv
Show abstract

Malawis HIV treatment monitoring system faces serious challenges because of a shortage of experts and reliance on viral load testing every 3 to 12 months. The process causes dangerous delays in identifying treatment failure. This leads to a higher risk of disease progression, transmission, and death. To tackle this issue, this study used a machine learning model based on association rules and combined it with clustering analysis to create a machine learning framework to identify key factors and risk profiles for virological failure among children living with HIV (CLHIV) in Malawi. The methodology combines a Random Forest classifier for feature importance, association rule mining to find predictive rules, and k-Prototype clustering for risk profiling among CLHIV. The random forest feature importance results show that Body Mass Index (BMI), CD4 count, TB status, ART regimen, gender, ART adherence, and treatment duration are major drivers of virological failure. In addition to these individual factors, the analysis produced highly reliable association rules with over 90% confidence. This establishes a framework for identifying complex risk profiles and informing focused clinical interventions. The high lift values of 4.9 across the most significant rules demonstrate the models effectiveness by revealing strong, non-random associations. Clustering analysis also identified two distinct risk profiles associated with virological failure. The k-prototype clustering model performed optimally with a cluster purity of 100% and a silhouette score of 79%.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.