Back

Utilizing large language models to construct a dataset of Württemberg's 19th-century fauna from historical records

Teich, M.; Escobari, B.; Rehbein, M.

2025-10-16 ecology
10.1101/2025.10.14.681982 bioRxiv
Show abstract

Constructing datasets on past biodiversity from historical sources is crucial for understanding long-term ecological changes. Typically, compiling such datasets relies on prior knowledge of the sources composition and requires considerable manual effort. To overcome these challenges, we implement an automated approach based on prompted large language models (LLMs) to detect mentions of species in texts from 19th-century Wurttemberg and link these mentions to identifiers in the GBIF database. Based on our evaluation, we find that LLMs can reliably identify species in the texts with high recall (92.6%) and precision (95.3%), while providing estimates of the correct species identifier with considerable accuracy (83.0%).

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.