Back

Methods for estimating location from spatial patterns in species composition: a fishing location case study

Smith, J. A.

2025-12-02 ecology
10.64898/2025.11.30.691468 bioRxiv
Show abstract

The location at which a species assemblage is observed can be uncertain, for example in fisheries where catch composition is recorded accurately but location data may be coarse or inaccurate. In such cases, spatial signals in observed species compositions can help identify and refine uncertain locations. This study explored three approaches for location estimation using species composition data: 1) a hierarchical species distribution model (SDM) that jointly estimates species distributions and location uncertainty, 2) an inverse prediction method that uses a fitted SDM to identify the most likely location given new species data, and 3) the direct modeling of location as the response variable. Each approach requires a subset of observations with accurate locations to quantify the spatial patterns in species distributions. All three methods were useful and reasonably accurate: for the simulated data the average distance error (distance to true location) was 15% the size of the domain, and for the real data this error was 20-100 km. When only some locations are uncertain, the hierarchical approach is valuable due to its integrated estimation of location and species parameters. When many locations are uncertain, the other approaches seem more suitable. Inverse prediction is ideal when prior information is available to constrain the location estimation. Direct modelling (although causally spurious) is well suited to large datasets, with multivariate random forests often providing the most accurate location estimates. These methods can enhance spatial data quality in ecological monitoring and help develop tools for improving inaccurate or deliberately misreported fishing locations.

Published in Ecological Modelling (predicted rank #13) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.