Back

Investigating Bacillus anthracis genomic diversity and trait-specific lineages in an endemic area in northern Tanzania through a combination of traditional and culture-free sequencing approaches

Hilbig, A.; Medvecky, M.; Lembo, T.; Mmbaga, B. T.; Kiwelu, I.; Mshanga, D.; Motto, S. K.; Makondo, Z. E.; Wadugu, B.; Arola, H. O.; Nikkari, S.; Rubach, M. P.; Crump, J. A.; Lycett, S. J.; Biek, R.; Forde, T. L.

2025-12-10 genomics
10.64898/2025.12.04.692280 bioRxiv
Show abstract

Anthrax, caused by Bacillus anthracis (BA), is a prominent neglected zoonosis with major impacts on human, livestock, and wildlife health. Despite this, limited genomic investigation at the One Health interface constrains current understanding of BA transmission and of the ecological and host factors shaping its diversity and population structure. This includes the possibility of host-specific BA lineages, given that anthrax outbreaks often disproportionally affect individual species. This study characterises the genomic diversity of BA in an endemic area, the Ngorongoro Conservation Area (NCA), in northern Tanzania. We analysed 213 BA genomes from livestock, wildlife and humans from cultured isolates combined with a culture-free targeted capture (TC) approach. NCA sequences formed a distinct genetic cluster compared with those from surrounding areas, and we observed surprisingly high levels of strain diversity within apparent epidemiological clusters, as well as within single animals, though strain diversity was lowest at the within host scale. We found limited evidence for seasonal clustering of cases as well as for BA lineages clustering by host species. This indicates that disproportional impacts on certain species during outbreaks are more likely driven by host ecology factors or, hypothetically, by accessory parts of the bacterial genome not represented in our data. TC-derived data significantly expanded the range of host species and geographic locations for genomic analysis, demonstrating the value of this approach. Although TC data may contain artefactual variation, shared SNP profiles between isolate- and TC-derived genomes gave confidence in its use for genotyping. Our analysis demonstrates unexpectedly high BA strain diversity and limited population structure in this endemic area across a range of spatial scales, including within-host. It further highlights the need for high density sampling and adaptable sequencing strategies to generate adequate BA genomic datasets that can enable informative molecular epidemiological studies of anthrax at the One Health interface. Author SummaryAnthrax continues to threaten the health of people, livestock, and wildlife in many parts of the world, yet we still know surprisingly little about how this disease spreads in nature. One major gap is understanding why some species are affected more strongly during outbreaks, despite assumed equal susceptibility. To investigate this, we studied the genomic diversity of the anthrax-causing bacterium Bacillus anthracis in a large conservation area in northern Tanzania. We combined two ways of generating genetic data: traditional laboratory culture and a culture-free method that allowed us to recover bacterial DNA from a wider range of samples. By analysing a uniquely large and species-diverse dataset of over 200 bacterial genomes from people, livestock and wildlife, we found that bacteria from the study area formed a clearly defined group compared to those from surrounding regions. We also discovered unexpectedly high diversity of strains, not only across the landscape but even within single animals. Despite this diversity, we saw little evidence that certain bacterial lineages are tied to specific host species. Our results suggest that the behaviour and ecology of different animals, rather than host-adapted lineages of the bacterium, likely explain why some species are more affected than others.

Published in PLOS Neglected Tropical Diseases (predicted rank #2) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.