Synonym Augmentation for Rare Disease Identification in Unstructured Data
Valinejad, J.; Moon, S.; Xu, Y.; Zhu, Q.
Show abstract
The significant challenges associated with rare diseases in the medical and research domains include the scarcity of information, which is often confined to unstructured formats. Although existing approaches provide valuable insights, there is a need to develop effective methods to identify information pertinent to rare diseases for advancing rare disease research. We identified mentions of rare diseases in relevant texts and assessed their relevance using derived scores, the confidence score and semantic similarity from a fine-tuned BioMedBERT encoder. This encoder was fine-tuned using rare disease related text from Online Mendelian Inheritance in Man (OMIM), Orphanet, a manually validated dataset, and STS benchmark datasets. The process of identifying meaningful rare disease mentioned was presented through two case studies that retrieved relevant NIH-funded projects, utilizing a generated knowledge graph in Neo4j to host data on 2,067 GARD diseases with over 320,000 NIH funded projects. Through various case studies with NIH-funded projects related to rare diseases, we demonstrated the effectiveness of our approach in systematically providing rare disease related data to enhance our understanding of rare diseases for future investigations.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DISEASES 2.0: a weekly updated database of disease-gene associations from text mining and data integration 96%
- A Sequence Labeling Framework for Extracting Drug-Protein Relations from Biomedical Literature 95%
- SynLethDB 2.0: A web-based knowledge graph database on synthetic lethality for novel anticancer drug discovery 95%
Similar papers in this journal
Similar papers in this journal
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 96%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 93%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.