Specifind: A Natural Language Processing Tool for Automating Species Occurrence (Re-)Discovery from Scientific Literature
Golomb Duran, T.; Far, J. A.; Diaz, A.; Barroso, M.; Roldan, A.; Colinas, N.; Cancellario, T.
Show abstract
A vast amount of valuable information on species occurrences remains embedded within the unstructured and continuously expanding body of ecological literature written in natural language. When effectively extracted, such dispersed knowledge has the potential to significantly improve understanding of species distributions and ecological patterns, thereby enabling more targeted and informed conservation actions. To address this challenge, Specifind is introduced as a tool designed to facilitate the extraction of species occurrence data through the identification of scientific species names, geographic references, and the relationships between them while ensuring traceability. At its core lies a newly developed and expertly annotated dataset comprising over one thousand open-access abstracts drawn from five domains: biogeography, botany, entomology, mycology, and zoology. A suite of Natural Language Processing components has been integrated, including Document Layout Analysis, Optical Character Recognition, Named Entity Recognition, Coreference Resolution, and Relation Extraction. These components enable the accurate identification and contextual linking of species and geographic entities, even when references are distributed throughout the text. As a result, Specifind enhances the discoverability and usability of species occurrence information embedded within unstructured scientific texts. It is expected to reduce the substantial effort required for manual literature review, thereby supporting biodiversity research, informing conservation planning, and facilitating ecological analysis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CoNECo: A Corpus for Named Entity recognition and normalization of protein Complexes 94%
- Understanding Ecological Systems Using Knowledge Graphs: An Application to Highly Pathogenic Avian Influenza 93%
- GRU-SCANET: Unleashing the Power of GRU-based Sinusoidal CApture Network for Precision-driven Named Entity Recognition 92%
Similar papers in this journal
Similar papers in this journal
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 92%
- Sequence Compression Benchmark (SCB) database - a comprehensive evaluation of reference-free compressors for FASTA-formatted sequences 91%
- Machado: open source genomics data integration framework 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.