Self-normalizing learning on biomedical ontologies using a deep Siamese neural network
Smaili, F. Z.; Gao, X.; Hoehndorf, R.
Show abstract
MotivationOntologies are widely used in biomedicine for the annotation and standardization of data. One of the main roles of ontologies is to provide structured background knowledge within a domain as well as a set of labels, synonyms, and definitions for the classes within a domain. The two types of information provided by ontologies have been extensively exploited in natural language processing and machine learning applications. However, they are commonly used separately, and thus it is unknown if joining the two sources of information can further benefit data analysis tasks. ResultsWe developed a novel method that applies named entity recognition and normalization methods on texts to connect the structured information in biomedical ontologies with the information contained in natural language. We apply this normalization both to literature and to the natural language information contained within ontologies themselves. The normalized ontologies and text are then used to generate embeddings, and relations between entities are predicted using a deep Siamese neural network model that takes these embeddings as input. We demonstrate that our novel embedding and prediction method using self-normalized biomedical ontologies significantly outperforms the state-of-the-art methods in embedding ontologies on two benchmark tasks: prediction of interactions between proteins and prediction of gene-disease associations. Our method also allows us to apply ontology-based annotations and axioms to the prediction of toxicological effects of chemicals where our method shows superior performance. Our method is generic and can be applied in scenarios where ontologies consisting of both structured information and natural language labels or synonyms are used. Availabilityhttps://github.com/bio-ontology-research-group/Ontology-based-normalization Contactrobert.hoehndorf@kaust.edu.sa and xin.gao@kaust.edu.sa
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CoNECo: A Corpus for Named Entity recognition and normalization of protein Complexes 95%
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 95%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 94%
Similar papers in this journal
- RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature 95%
- A Sequence Labeling Framework for Extracting Drug-Protein Relations from Biomedical Literature 95%
- SynLethDB 2.0: A web-based knowledge graph database on synthetic lethality for novel anticancer drug discovery 94%
Similar papers in this journal
- HARVESTMAN: A framework for hierarchical featurelearning and selection from whole genome sequencingdata 94%
- SKiM-GPT: Combining Biomedical Literature-Based Discovery with Large Language Model Hypothesis Evaluation 94%
- Optimizing biomedical information retrieval with a keyword frequency-driven Prompt Enhancement Strategy 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.