Back

Deep Multi-modal Species Occupancy Modeling

Haucke, T.; Harrell, L.; Shen, Y.; Klein, L.; Rolnick, D.; Gillespie, L. E.; Beery, S.

2025-09-11 ecology
10.1101/2025.09.06.674602 bioRxiv
Show abstract

Effective conservation and restoration of species is an increasingly urgent priority. To design management strategies that improve species success, we need a solid understanding of the habitat characteristics that support it. Occupancy models are statistical tools that ecologists use to model these relationships from data. Yet, current models represent habitats with coarse-scale environmental variables that fail to capture important microhabitat features. We show that these limitations can be addressed by incorporating AI-derived, multimodal habitat representations from overhead satellite imagery and ground-level camera-trap imagery. Across geography and species, these representations yield more accurate out-of-sample predictions than models based on conventional covariates alone, and combining satellite and ground-level views provides complementary gains. To translate improved prediction into actionable ecological insight, we further introduce a method that makes black-box AI-derived habitat representations interpretable by summarizing key factors contributing to occupancy probability into text-based descriptions. We then generate a per-site score for each description, which can replace black-box features to transparently link discovered habitat elements to species occurrence while maintaining predictive performance. Our approach provides a path toward microhabitat-aware and interpretable species-habitat models that support restoration planning and management decisions. We implement our method in an open-source Python package bridging AI and statistical ecology.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.