Back

Global vascular plants reveal persistent gaps across taxa and ecoregions

Maciel, E. A.

2026-08-28 evolutionary biology
10.64898/2026.08.24.746674 bioRxiv
Show abstract

Biodiversity aggregators such as GBIF provide unprecedented access to global biodiversity data, yet their representativeness remains uneven across space and taxa. This study examined the spatial and taxonomic structure of global vascular plant data available on GBIF. Six filters were applied to the GBIF vascular plant dataset, resulting in the removal of 54% of all records. Together, the filters explained more than 90% of the identified spatial issues, with duplicate and missing coordinates accounting for most of the variation. A higher number of occurrence records was associated with a greater number of spatial issues. Record distributions became progressively more even at finer taxonomic levels, from orders to species. The time series of occurrences for species, genera, and families increased sharply after 1800 and continued to rise, with no apparent stabilisation. Of the 824 ecoregions covered, 73 accounted for 72% of all occurrence records. These ecoregions spanned all continents but were strongly concentrated in Europe, followed by North America and Oceania. The analyses reveal four key patterns: (1) data volume is positively associated with spatial issues; (2) a small number of taxa account for a large proportion of records, whereas many are represented by relatively few; (3) occurrence data aggregated by GBIF have increased continuously since 1800; and (4) record coverage remains highly uneven across the world's ecoregions. These results highlight the substantial contribution of biodiversity data aggregators to expanding access to biological information while demonstrating the persistent spatial and taxonomic biases that shape their contents. Such biases should be explicitly considered when assessing data completeness and quality and when using aggregated occurrence records to infer global biodiversity patterns.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.