MicrobioRel: A Set of Datasets for Microbiome Relation Extraction
EL KHETTARI, O.; Batteux, D.; Quiniou, S.; Chaffron, S.
Show abstract
Biomedical knowledge curation relies on a variety of Natural Language Processing tasks, including biomedical entity recognition and document-level relation extraction. With the growing size and capabilities of Language Models, effectively deploying them in specific and specialised domains remains a persistent challenge, highlighting the need for high-quality, domain-adapted datasets. In this work, we present MicrobioRel, a corpus of two datasets to study the relations between biological entities in the gut microbiome. The first dataset, MicrobioRel-cur, is a document-level labelled corpus, corresponding to paragraphs from journal articles that were manually annotated with different types of relations between biomedical concepts in the gut microbiome domain. We describe its creation process, annotation guidelines, and key statistics. On this dataset, we evaluated different architectures for relation extraction and identify PubMedBERT as the most effective model for this task. We also created a second dataset, MicrobioRel-pred, by generating relation predictions on other journal articles using the fine-tuned PubMedBERT model. We demonstrate its potential to extract meaningful interactions. The MicrobioRel is a crucial resource for advancing tasks like automatic knowledge extraction in specialised domains such as the gut microbiome, facilitating hypothesis generation and supporting scientific discovery.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- LSD600: the first corpus of biomedical abstracts annotated with lifestyle–disease relations 96%
- RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature 95%
- DISEASES 2.0: a weekly updated database of disease-gene associations from text mining and data integration 94%
Similar papers in this journal
Similar papers in this journal
- Knowledge Graph-based Thought: a knowledge graph enhanced LLMs framework for pan-cancer question answering 94%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 94%
- Strategies and Techniques for Quality Control and Semantic Enrichment with Multimodal Data: A Case Study in Colorectal Cancer with eHDPrep 94%
Similar papers in this journal
- ProtAlign-ARG: Antibiotic Resistance Gene Characterization Integrating Protein Language Models and Alignment-Based Scoring 93%
- A natural language processing system for the efficient extraction of cell markers 93%
- Hypothesizing mechanistic links between microbes and disease using knowledge graphs 93%
Similar papers in this journal
- Alzheimer Disease Knowledge Graph Enhances Knowledge Discovery and Disease Prediction 94%
- A fast, accurate, and generalisable heuristic-based negation detection algorithm for clinical text 93%
- Inference of disease-associated microbial gene modules based on metagenomic and metatranscriptomic data 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.