Species distribution modeling for disease ecology: a multi-scale case study for schistosomiasis host snails in Brazil
Singleton, A. L.; Glidden, C. K.; Chamberlin, A. J.; Tuan, R.; Palasio, R. G. S.; Pinter, A.; Caldeira, R. L.; Mendonca, C. L. F.; Carvalho, O. S.; Monteiro, M. V.; Athni, T. S.; Sokolow, S. H.; Mordecai, E. A.; De Leo, G. A.
Show abstract
Species distribution models (SDMs) are increasingly popular tools for profiling disease risk in ecology, particularly for infectious diseases of public health importance that include an obligate non-human host in their transmission cycle. SDMs can create high-resolution maps of host distribution across geographical scales, reflecting baseline risk of disease. However, as SDM computational methods have rapidly expanded, there are many outstanding methodological questions. Here we address key questions about SDM application, using schistosomiasis risk in Brazil as a case study. Schistosomiasis--a debilitating parasitic disease of poverty affecting over 200 million people across Africa, Asia, and South America--is transmitted to humans through contact with the free-living infectious stage of Schistosoma spp. parasites released from freshwater snails, the parasites obligate intermediate hosts. In this study, we compared snail SDM performance across machine learning (ML) approaches (MaxEnt, Random Forest, and Boosted Regression Trees), geographic extents (national, regional, and state), types of presence data (expert-collected and publicly-available), and snail species (Biomphalaria glabrata, B. tenagophila and B. straminea). We used high-resolution (1km) climate, hydrology, land-use/land-cover (LULC), and soil property data to describe the snails ecological niche and evaluated models on multiple criteria. Although all ML approaches produced comparable spatially cross-validated performance metrics, their suitability maps showed major qualitative differences that required validation based on local expert knowledge. Additionally, our findings revealed varying importance of LULC and bioclimatic variables for different snail species at different spatial scales. Finally, we found that models using publicly-available data predicted snail distribution with comparable AUC values to models using expert-collected data. This work serves as an instructional guide to SDM methods that can be applied to a range of vector-borne and zoonotic diseases. In addition, it advances our understanding of the relevant environment and bioclimatic determinants of schistosomiasis risk in Brazil.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- An ecological niche model to predict the geographic distribution of Haemagogus janthinomys, Dyar, 1921 the yellow fever and Mayaro virus vector, in South America 95%
- Modelling geospatial distributions of the triatomine vectors of Trypanosoma cruzi in Latin America 93%
- The impact of climate suitability, urbanisation, and connectivity on the expansion of dengue in 21st century Brazil 93%
Similar papers in this journal
- Field-validation of multiple species distribution models shows variation in performance for predicting Aedes albopictus distributions at the invasion edge 93%
- Diversity and seasonality of ectoparasite burden on two species of Madagascar fruit bat, Eidolon dupreanum and Rousettus madagascariensis 92%
- Prediction of mosquito vector abundance for three species in the Anopheles gambiae complex 92%
Similar papers in this journal
Similar papers in this journal
- EpiFusion: Joint inference of the effective reproduction number by integrating phylodynamic and epidemiological modelling with particle filtering 92%
- Call detail record aggregation methodology impacts infectious disease models informed by human mobility 91%
- Spatial distribution of poultry farms using point pattern modelling: a method to address livestock environmental impacts and disease transmission risks 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.