Back

Protocol for the production of an Arenavirus and Hantavirus host-pathogen database: Project ArHa.

Simons, D.; Rivero, R.; Martinez-Checa Guiote, A.; Gordon, H.; Milne, G. C.; Rickard, G.; Redding, D. W.; Seifert, S. N.

2025-01-18 microbiology
10.1101/2025.01.17.633514 bioRxiv
Show abstract

1Arenaviruses and Hantaviruses, primarily hosted by rodents and shrews, represent significant public health threats due to their potential for zoonotic spillover into human populations. Despite their global distribution, the full impact of these viruses on human health remains poorly understood, particularly in regions like Africa, where data is sparse. Both virus families continue to emerge, with pathogen evolution and spillover driven by anthropogenic factors such as land use change, climate change, and biodiversity loss. Recent research highlights the complex interactions between ecological dynamics, host species, and environmental factors in shaping the risk of pathogen transmission and spillover. This underscores the need for integrated ecological and genomic approaches to better understand these zoonotic diseases. A comprehensive, spatially and temporally explicit dataset, incorporating host-pathogen dynamics and human disease data, is crucial for improving risk assessments, enhancing disease surveillance, and guiding public health interventions. Such a dataset (ArHa) would also support predictive modelling efforts aimed at mitigating future spillover events. This paper proposes the development of this unified database for small-mammal hosts of Arenaviruses and Hantaviruses, identifying gaps in current research and promoting a more comprehensive understanding of pathogen prevalence, spillover risk, and viral evolution. 2 Strengths and Limitations of this studyO_LIThis dataset combines detailed spatial and temporal information, providing a unique resource for understanding geographic and temporal trends in Arenavirus and Hantavirus host-pathogen relationships. C_LIO_LIBy explicitly quantifying sampling biases and detection efforts, the dataset allows more robust and accurate asssessments of pathogen prevalence and distribution. C_LIO_LIThe dataset offers a platform for linking ecological data with human health outcomes, supporting the identification of spillover hotspots. C_LIO_LIThe dataset relies on published material, which may vary in terms of detail, accuracy and completeness. Missing or imprecise information may limit the reliability of subsequent analyses. C_LIO_LIThe dataset will be produced as a static resource which could limit its relevance over time as emerging data will not be added. C_LI

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.