Back

Network Analysis of Pairwise Relative Tuberculosis Transmission Probabilities in Lima, Peru

Shapiro, A. N.; Brooks, M. B.; Huang, C.; Murray, M. B.; White, L. F.; Jenkins, H. E.

2025-11-19 infectious diseases
10.1101/2025.11.18.25340467 medRxiv
Show abstract

BackgroundIdentifying transmission events is important in understanding infectious disease dynamics. Such events are typically unobservable, particularly in diseases with long serial intervals such as tuberculosis (TB). We apply network techniques to identify transmission clusters and features shared within clusters. MethodsWe estimate directed pairwise transmission probabilities via an existing iterative algorithm that employs a modified Naive Bayes classifier to incorporate demographic, clinical, and genetic data and use these probabilities to create a network. We explore noise reduction techniques to trim low probability edges. We apply clustering algorithms to group together individuals with TB based on edges informed by transmission probabilities. We apply our framework to simulated data and assess how the clustering algorithms captured the simulated clusters. We then apply this approach to data from a cohort study in Lima, Peru and examine the homogeneity of the clusters using a binary entropy measure. ResultsWe find cluster performance to be consistent across all edge trimming scenarios and clustering methods. We find high levels of entropy for age, sex, socioeconomic status, and individuals who work outside the house and use public transit, indicating these variables are heterogenous across clusters. ConclusionsWe demonstrate approaches to analyze estimated directed pairwise transmission probabilities with network techniques. The approach is consistent across network construction and clustering methods. This method can be applied to any disease outbreak to understand its dynamics.

Published in American Journal of Epidemiology (predicted rank #6) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.