Back

biodumpy: A Comprehensive Biological Data Downloader

Cancellario, T.; Golomb Duran, T.; Far Morenilla, A. J.; Roldan, A.; Capa, M.

2025-07-29 ecology
10.1101/2025.07.26.666724 bioRxiv
Show abstract

In recent years, the expansion of public biodiversity platforms and associated datasets has greatly improved access to ecological and biological information. These resources now cover vast geographic areas, extended temporal scales, and diverse taxonomic groups, becoming essential for ecological studies by enabling more comprehensive analyses and novel hypotheses testing. Concurrently, the development of programming packages has facilitated data access and interaction, streamlining their retrieval processes. However, most existing tools are limited to specific databases, posing challenges for studies requiring seamless integration of data from multiple sources. The growing availability of biodiversity data highlights the urgent need for robust tools to efficiently process, analyse, and interpret ecological and biological information. To address this limitation, we introduce biodumpy, a new Python package developed for the retrieval, management, and integration of biological data from various public databases. biodumpy provides access to up-to-date and comprehensive datasets spanning genetic, distributional, taxonomic, and bibliographic sources. It includes specialized modules for efficient data retrieval across taxonomic lists, with the capability to process multiple modules simultaneously. By integrating diverse data sources, biodumpy enhances data acquisition, providing researchers with a powerful framework for comprehensive analyses and supporting ecological research to tackle complex environmental challenges.

Published in Biodiversity Informatics · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.