Back

Improving estimation of species distribution from citizen-science records using data-integration models

Zulian, V.; Miller, D. A. W.; Ferraz, G.

2021-04-11 ecology
10.1101/2021.04.09.439158 bioRxiv
Show abstract

Mapping species distributions is a crucial but challenging requirement of wildlife management. The frequent need to sample vast expanses of potential habitat increases the cost of planned surveys and rewards accumulation of opportunistic observations. In this paper, we integrate planned survey data from roost counts with opportunistic samples from eBird, WikiAves and Xeno-canto citizen-science platforms to map the geographic range of the endangered Vinaceous-breasted Parrot. We demonstrate the estimation and mapping of species occurrence based on data integration while accounting for specifics of each data set, including observation technique and uncertainty about the observations. Our analysis illustrates 1) the incorporation of sampling effort, spatial autocorrelation, and site covariates in a joint-likelihood, hierarchical, data-integration model; 2) the evaluation of the contribution of each data set, as well as the contribution of effort covariates, spatial autocorrelation, and site covariates to the predictive ability of fitted models using a cross-validation approach; and 3) how spatial representation of the latent occupancy state (i.e. realized occupancy) helps identify areas with high uncertainty that should be prioritized in future field work. Our results reveal a Vinaceous-breasted Parrot geographic range of 434,670 km2, which is three times larger than the Extant area previously reported in the IUCN Red List. The exclusion of one data set at a time from the analyses always resulted in worse predictions by the models of truncated data than by the full model, which included all data sets. Likewise, exclusion of spatial autocorrelation, site covariates, or sampling effort resulted in worse predictions. The integration of different data sets into one joint-likelihood model produced a more reliable representation of the species range than any individual data set taken on its own improving the use of citizen science data in combination with planned survey results.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.