Exploring binary relations for ontology extension and improved adaptation to clinical text
Slater, L. T.; Hoehndorf, R.; Karwath, A.; Gkoutos, G. V.
Show abstract
BackgroundThe controlled domain vocabularies provided by ontologies make them an indispensable tool for text mining. Ontologies also include semantic features in the form of taxonomy and axioms, which make annotated entities in text corpora useful for semantic analysis. Extending those semantic features may improve performance for characterisation and analytic tasks. Ontology learning techniques have previously been explored for novel ontology construction from text, though most recent approaches have focused on literature, with applications in information retrieval or human interaction tasks. We hypothesise that extension of existing ontologies using information mined from clinical narrative text may help to adapt those ontologies such that they better characterise those texts, and lead to improved classification performance. ResultsWe develop and present a framework for identifying new classes in text corpora, which can be integrated into existing ontology hierarchies. To do this, we employ the Stanford Open Information Extraction algorithm and integrate its implementation into the Komenti semantic text mining framework. To identify whether our approach leads to better characterisation of text, we present a case study, using the method to learn an adaptation to the Disease Ontology using text associated with a sample of 1,000 patient visits from the MIMIC-III critical care database. We use the adapted ontology to annotate and classify shared first diagnosis on patient visits with semantic similarity, revealing an improved performance over use of the base Disease Ontology on the set of visits the ontology was constructed from. Moreover, we show that the adapted ontology also improved performance for the same task over two additional unseen samples of 1,000 and 2,500 patient visits. ConclusionsWe report a promising new method for ontology learning and extension from text. We demonstrate that we can successfully use the method to adapt an existing ontology to a textual dataset, improving its ability to characterise the dataset, and leading to improved analytic performance, even on unseen portions of the dataset.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 98%
- Towards semantic interoperability: finding and repairing hidden contradictions in biomedical ontologies 97%
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 95%
Similar papers in this journal
- Extract, Transform, Load Framework for the Conversion of Health Databases to OMOP 94%
- Understanding signaling and metabolic paths using semantified and harmonized information about biological interactions 94%
- Knowledge Beacons: Web Service Workflow for FAIR Data Harvesting of Distributed Biomedical Knowledge 94%
Similar papers in this journal
- The Xenopus Phenotype Ontology: bridging model organism phenotype data to human health and development. 95%
- RTX-KG2: a system for building a semantically standardized knowledge graph for translational biomedicine 95%
- Optimizing biomedical information retrieval with a keyword frequency-driven Prompt Enhancement Strategy 95%
Similar papers in this journal
- De-novo FAIRification via an Electronic Data Capture system by automated transformation of filled electronic Case Report Forms into machine-readable data 95%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 94%
- Medication information extraction using local large language models 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.