Clustering rare diseases within an ontology-enriched knowledge graph
Sanjak, J.; Mathe, E.; Zhu, Q.
Show abstract
Structured AbstractO_ST_ABSObjectiveC_ST_ABSIdentifying sets of rare diseases with shared aspects of etiology and pathophysiology may enable drug repurposing and/or platform based therapeutic development. Toward that aim, we utilized an integrative knowledge graph-based approach to constructing clusters of rare diseases. Materials and MethodsData on 3,242 rare diseases were extracted from the National Center for Advancing Translational Science (NCATS) Genetic and Rare Diseases Information center (GARD) internal data resources. The rare disease data was enriched with additional biomedical data, including gene and phenotype ontologies, biological pathway data and small molecule-target activity data, to create a knowledge graph (KG). Node embeddings were used to convert nodes into vectors upon which k-means clustering was applied. We validated the disease clusters through semantic similarity and feature enrichment analysis. ResultsA node embedding model was trained on the ontology enriched rare disease KG and k-means clustering was applied to the embedding vectors resulting in 37 disease clusters with a mean size of 87 diseases. We validate the disease clusters quantitatively by looking at semantic similarity of clustered diseases, using the Orphanet Rare Disease Ontology. In addition, the clusters were analyzed for enrichment of associated genes, revealing that the enriched genes within clusters were shown to be highly related. DiscussionWe demonstrate that node embeddings are an effective method for clustering diseases within a heterogenous KG. Semantically similar diseases and relevant enriched genes have been uncovered within the clusters. Connections between disease clusters and approved or investigational drugs are enumerated for follow-up efforts. ConclusionOur study lays out a method for clustering rare diseases using the graph node embeddings. We develop an easy to maintain pipeline that can be updated when new data on rare diseases emerges. The embeddings themselves can be paired with other representation learning methods for other data types, such as drugs, to address other predictive modeling problems. Detailed subnetwork analysis and in-depth review of individual clusters may lead to translatable findings. Future work will focus on incorporation of additional data sources, with a particular focus on common disease data.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- HELP: A computational framework for labelling and predicting human common and context-specific essential genes 94%
- MENDELSEEK: An algorithm that predicts Mendelian Genes and elucidates what makes them special 94%
- Causal reasoning over knowledge graphs leveraging drug-perturbed and disease-specific transcriptomic signatures for drug discovery 94%
Similar papers in this journal
Similar papers in this journal
- Mining hidden knowledge: Embedding models of cause-effect relationships curated from the biomedical literature 94%
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 94%
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 94%
Similar papers in this journal
- Replacing non-biomedical concepts improves embedding of biomedical concepts 94%
- Combining explainable machine learning, demographic and multi-omic data to identify precision medicine strategies for inflammatory bowel disease 93%
- ProteinWeaver: A Webtool to Visualize Ontology-Annotated Protein Networks 93%
Similar papers in this journal
- Network and pathway expansion of genetic disease associations identifies successful drug targets 94%
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 94%
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.