Back

Understanding and addressing under-sampling in plant-pollinator networks using stacked models for missing link prediction

Van Kleunen, L. B.; de Manincor, N.; Dee, L. E.; Clauset, A.; Roberts, S.; Massol, F.

2026-01-09 ecology
10.64898/2026.01.08.698368 bioRxiv
Show abstract

Pollinator diversity and pollination are under threat from anthropogenic disturbances. Plant-pollinator networks are useful tools to study the consequences of these disturbances, but their links are often under-sampled, potentially biasing conclusions. We introduce a computational method that can be used to predict which unobserved links in a plant-pollinator network dataset are missing. This method extends a state-of-the-art machine learning approach to missing link prediction based on stacked generalization, which has been promising in other contexts, to the setting of bipartite plant-pollinator networks with species traits treated as node attributes. This approach predicts missing links using an ensemble of predictors based on both observed network topology and species traits. We demonstrate this approach on synthetic and empirical plant-pollinator networks. We take advantage of a unique empirical plant-pollinator dataset which samples from three sites using (a) records of observed visits of pollinators to plant species and (b) pollen analysis with detailed species trait annotations for both plant and pollinator nodes. Visit-based sampling under-samples the links in these networks, whereas pollen-based sampling identifies additional links, allowing us to investigate missing link prediction from realistic patterns of field under-sampling. We show that full plant-pollinator networks can be partially reconstructed from the under-sampled networks using our approach. Further, we show that this method can be used to identify which features are the most important for predicting missing links, which allows for investigation into mechanisms driving plant-pollinator link formation and observation. This method is broadly applicable to under-sampled plant-pollinator networks.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.