Back

OntoContext, a new python package for gene contextualization based on the annotation of biomedical texts

Bedhiafi, W.; Thomas-Vaslin, V.; Benammar Elgaaied, A.; Six, A.

2022-05-28 bioinformatics
10.1101/2022.05.27.493696 bioRxiv
Show abstract

MotivationThe automatic mining for bibliography exploitation in given contexts is a challenge according to the increasing number of scientific publications and new concepts. Several indexing systems were developed for biomedical literature. However, such systems have failed to produce contextualised research of genes and proteins and automatically group texts according to shared concepts. In this paper, we present OntoContext, a contextualization system crossing the use of biomedical ontologies to annotate texts containing terms related to cell populations, anatomical locations and diseases and to extract gene, RNA or protein names in these contexts. ResultsOntoContext, a new python package contains two modules. The "annot" module for "annotation" function, is based on combination of morphosyntactic labelling and exact matching and on dictionaries derived from the Cell Ontology, the UBERON Ontology (anatomical context), the Human Disease Ontology and geniatagger, (which contains particular tags for gene-related names). The "annot" output is used as input for the second module "crisscross" generating lists of gene-related names obtained by crossing annotations from the three mentioned ontologies. OntoContext showed better performances than NCBO Annotator after evaluation on two text corpuses. OntoContext is freely available in the pypi. Availabilityhttps://pypi.python.org/pypi/OntoContext and https://github.com/walidbedhiafi/OntoContext1. Contactadrien.six@sorbonne-universite.fr

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.