Back

Extracting massive ecological data on state and interactions of species using large language models

Keck, F.; Broadbent, H.; Altermatt, F.

2025-01-27 ecology
10.1101/2025.01.24.634685 bioRxiv
Show abstract

The contemporary ecological crisis calls for integration and synthesis of ecological data describing the state, change and processes of ecological communities. However, such synthesis depends on the integration of vast amounts of mostly scattered and often hard-to-extract information that is published and dispersed across hundreds of thousands of scientific papers, for example describing species-specific interactions and trophic relationships. Recent advancements in natural language processing (NLP) and in particular the emergence of large language models (LLMs) offer a novel, and potentially revolutionary solution to this persistent challenge, for the first time creating the opportunity to access and extract virtually all data ever published. Here, we demonstrate the transformative potential of LLMs by extracting all types of biological interactions among species directly from a corpus of 83,910 scientific articles. Our approach successfully extracted a network of 144,402 interactions between 36,471 taxa. Performance analysis shows that the model exhibits a high sensitivity (70.0%) and excellent precision (89.5%). Our approach proves that LLMs are capable of carrying out complex extraction tasks on key ecological data on a very large scale, paving the way for a multitude of potential applications in ecology and beyond.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.