Utilizing large language models to construct a dataset of Württemberg's 19th-century fauna from historical records
Teich, M.; Escobari, B.; Rehbein, M.
Show abstract
Constructing datasets on past biodiversity from historical sources is crucial for understanding long-term ecological changes. Typically, compiling such datasets relies on prior knowledge of the sources composition and requires considerable manual effort. To overcome these challenges, we implement an automated approach based on prompted large language models (LLMs) to detect mentions of species in texts from 19th-century Wurttemberg and link these mentions to identifiers in the GBIF database. Based on our evaluation, we find that LLMs can reliably identify species in the texts with high recall (92.6%) and precision (95.3%), while providing estimates of the correct species identifier with considerable accuracy (83.0%).
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Data-centric AI approach for automated wildflower monitoring 93%
- Geographic Name Resolution Service: A tool for the standardization and indexing of world political division names, with applications to species distribution modeling 93%
- Utilizing CNNs for classification and uncertainty quantification for 15 families of European fly pollinators 93%
Similar papers in this journal
- Large language models overcome the challenges of unstructured text data in ecology 96%
- The Soil Food Web Ontology: aligning trophic groups, processes, resources, and dietary traits to support food-web research 92%
- Hierarchical Classification of Insects with Multitask Learning and Anomaly Detection 92%
Similar papers in this journal
Similar papers in this journal
- Quantitative monitoring of nucleotide sequence data from genetic resources in context of their citation in the scientific literature 94%
- ExTaxsI: an exploration tool of biodiversity molecular data 93%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.