OntoContext, a new python package for gene contextualization based on the annotation of biomedical texts
Bedhiafi, W.; Thomas-Vaslin, V.; Benammar Elgaaied, A.; Six, A.
Show abstract
MotivationThe automatic mining for bibliography exploitation in given contexts is a challenge according to the increasing number of scientific publications and new concepts. Several indexing systems were developed for biomedical literature. However, such systems have failed to produce contextualised research of genes and proteins and automatically group texts according to shared concepts. In this paper, we present OntoContext, a contextualization system crossing the use of biomedical ontologies to annotate texts containing terms related to cell populations, anatomical locations and diseases and to extract gene, RNA or protein names in these contexts. ResultsOntoContext, a new python package contains two modules. The "annot" module for "annotation" function, is based on combination of morphosyntactic labelling and exact matching and on dictionaries derived from the Cell Ontology, the UBERON Ontology (anatomical context), the Human Disease Ontology and geniatagger, (which contains particular tags for gene-related names). The "annot" output is used as input for the second module "crisscross" generating lists of gene-related names obtained by crossing annotations from the three mentioned ontologies. OntoContext showed better performances than NCBO Annotator after evaluation on two text corpuses. OntoContext is freely available in the pypi. Availabilityhttps://pypi.python.org/pypi/OntoContext and https://github.com/walidbedhiafi/OntoContext1. Contactadrien.six@sorbonne-universite.fr
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature 94%
- DISEASES 2.0: a weekly updated database of disease-gene associations from text mining and data integration 93%
- LSD600: the first corpus of biomedical abstracts annotated with lifestyle–disease relations 93%
Similar papers in this journal
- The Xenopus Phenotype Ontology: bridging model organism phenotype data to human health and development. 94%
- SKiM-GPT: Combining Biomedical Literature-Based Discovery with Large Language Model Hypothesis Evaluation 92%
- Optimizing biomedical information retrieval with a keyword frequency-driven Prompt Enhancement Strategy 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.