Back

HIV-phyloTSI: Subtype-independent estimation of time since HIV-1 infection for cross-sectional measures of population incidence using deep sequence data

Golubchik, T.; Abeler-Dorner, L.; Hall, M.; Wymant, C.; Bonsall, D.; Macintyre-Cockett, G.; Thomson, L.; Baeten, J. M.; Celum, C. L.; Galiwango, R. M.; Kosloff, B.; Limbada, M.; Mujugira, A.; Mugo, N. R.; Gall, A.; Blanquart, F.; Bakker, M.; Bezemer, D.; Ong, S. H.; Albert, J.; Bannert, N.; Fellay, J.; Gunsenheimer-Bartmeyer, B.; Gunthard, H. F.; Kivela, P.; Kouyos, R. D.; Meyer, L.; Porter, K.; van Sighem, A.; van der Valk, M.; Berkhout, B.; Kellam, P.; Cornelissen, M.; Reiss, P.; Ayles, H.; Burns, D. N.; Fidler, S.; Grabowski, M. K.; Hayes, R.; Herbeck, J. T.; Kagaayi, J.; Kaleebu, P.; Linga

2022-05-16 hiv aids
10.1101/2022.05.15.22275117 medRxiv
Show abstract

Estimating the time since HIV infection (TSI) at population level is essential for tracking changes in the global HIV epidemic. Most methods for determining duration of infection classify samples into recent and non-recent and are unable to give more granular TSI estimates. These binary classifications have a limited recency time window of several months, therefore requiring large sample sizes, and cannot assess the cumulative impact of an intervention. We developed a Random Forest Regression model, HIV-phyloTSI, that combines measures of within-host diversity and divergence to generate TSI estimates from viral deep-sequencing data, with no need for additional variables. HIV-phyloTSI provides a continuous measure of TSI up to 9 years, with a mean absolute error of less than 12 months overall and less than 5 months for infections with a TSI of up to a year. It performed equally well for all major HIV subtypes based on data from African and European cohorts. We demonstrate how HIV-phyloTSI can be used for incidence estimates on a population level.

Published in BMC Bioinformatics (predicted rank #16) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.