Back

BORD: A Biomedical Ontology based method for concept Recognition using Distant supervision: Application to Phenotypes and Diseases

Toonsi, S.; Kafkas, S.; Hoehndorf, R.

2023-02-16 bioinformatics
10.1101/2023.02.15.528695 bioRxiv
Show abstract

MotivationConcept recognition in biomedical text is an important yet challenging task. The two main approaches to recognize concepts in text are dictionary-based approaches and supervised machine learning approaches. While dictionary-based approaches fail in recognising new concepts and variations of existing concepts, supervised methods require sufficiently large annotated datasets which are expensive to obtain. Methods based on distant supervision have been developed to use machine learning without large annotated corpora. However, for biomedical concept recognition, these approaches do not yet exploit the context in which a concept occurs in literature, and they do not make use of prior knowledge about dependencies between concepts. ResultsWe developed BORD, a Biomedical Ontology-based method for concept Recognition using Distant supervision. BORD utilises context from corpora which are lexically annotated using labels and synonyms from the classes of a biomedical ontology for model training. Furthermore, BORD utilises the ontology hierarchy for normalising the recognised mentions to their concept identifiers. We show how our method improves the performance of state of the art methods for recognising disease and phenotype concepts in biomedical literature. Our method is generic, does not require manually annotated corpora, and is robust to identify mentions of ontology classes in text. Moreover, to the best of our knowledge, this is the first approach utilising the ontology hierarchy for concept recognition. AvailabilityBORD is publicly available from https://github.com/bio-ontology-research-group/BORD Contactrobert.hoehndorf@kaust.edu.sa

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.