Drug repurposing for rare diseases via a gene-bridged heterogeneous knowledge graph and graph attention network
Ramani, D.
Show abstract
Rare diseases are severely underserved by pharmacological treatments, and computational drug repurposing offers a cost-effective alternative to de novo discovery. We present a reproducible end-to-end pipeline integrating 3,961 rare disease-gene associations from Orphadata with 98,239 gene-drug records from DisGeNET through a multi-stage harmonization pipeline (HGNC symbol standardization and RapidFuzz fuzzy matching), yielding a large-scale gene-bridged rare disease tripartite knowledge graph -- to our knowledge the largest such graph constructed exclusively from Orphadata and DisGeNET-- comprising 15,454 nodes and 35,131 edges spanning 2,249 clinically distinct rare diseases. A Graph Attention Network (GAT) trained on node-type classification as a pretext task achieves macro F1 = 0.651 and ROC-AUC = 0.818 on a stratified held-out test set, with stable performance across five evaluation partitions (SD [≤] 0.007). Drug candidate retrieval via cosine similarity in the GAT embedding space achieves Hits@10 = 0.400 across 200 evaluated disorders (vs. < 0.001 random baseline), with the clinically validated drug NITISINONE recovered at rank 4 for a tyrosine catabolism pathway disorder without pathway annotations. A deployment-ready interface is publicly available on HuggingFace Spaces.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Graph-Attention-Based Deep Learning Network for Predicting Biotech-Small-Molecule Drug Interactions 94%
- Prompt-to-Pill: Multi-Agent Drug Discovery and Clinical Simulation Pipeline 93%
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 93%
Similar papers in this journal
Similar papers in this journal
- Towards explainable interaction prediction: Embedding biological hierarchies into hyperbolic interaction space 93%
- Two-step multi-omics modelling of drug sensitivity in cancer cell lines to identify driving mechanisms 93%
- Predicting compound-protein interaction using hierarchical graph convolutional networks 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.