STEM-LM: Spatio-Temporal Ecological Modeling via Masked Language Model for Joint Species Distribution
Li, J. K.; Lim, W.; Callahan, F. M.; Raskin, L. Y.; Lemmon-Kishi, M.; Nielsen, R.
Show abstract
Joint species distribution models (JSDMs) are central to biodiversity forecasting and conservation decision-making. As ecological datasets grow in size, dimensionality, and spatio-temporal resolution, there is a need for flexible yet scalable JSDMs tailored to large-scale species observation data. Recent advances in masked language modeling for text and genomics suggest a natural alternative: by treating each species presence or absence as a token, and a sites species assemblage together with its spatio-temporal and ecological covariates as a sentence, we can learn joint co-occurrence structure by reconstructing masked species from their neighboring sites. We propose STEM-LM1, a Transformer-based JSDM that frames joint species distribution modeling as masked language modeling. By varying the masking rate during training, a single trained model supports both purely spatiotemporal/ecological prediction and conditioning on arbitrary subsets of observed species for joint co-occurrence inference at a given site. On a North American butterfly and a global plant distribution dataset, STEM-LM performs better or on par with other statistical and deep-learning based methods in terms of discriminative ranking, while producing substantially better rank-calibrated occurrence probabilities. Utilizing partial species observations at the same site greatly enhances prediction performance.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Common to rare transfer learning (CORAL) enables inference and prediction for a quarter million rare Malagasy arthropods 95%
- Towards Universal Cell Embeddings: Integrating Single-cell RNA-seq Datasets across Species with SATURN 94%
- Discovering novel cell types across heterogeneous single-cell experiments 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.