Path-based reasoning in biomedical knowledge graphs with BioPathNet
Hu, Y.; Oleshko, S.; Firmani, S.; Zhu, Z.; Cheng, H.; Ulmer, M.; Arnold, M.; Colome-Tatche, M.; Tang, J.; Xhonneux, S.; Marsico, A.
Show abstract
Understanding complex interactions in biomedical networks is crucial for advancements in biomedicine, but traditional link prediction (LP) methods are limited in capturing this complexity. Representation-based learning techniques improve prediction accuracy by mapping nodes to low-dimensional embeddings, yet they often struggle with interpretability and scalability. We present BioPathNet, a novel graph neural network framework based on the Neural Bellman-Ford Network (NBFNet), addressing these limitations through path-based reasoning for LP in biomedical knowledge graphs. Unlike node-embedding frameworks, BioPathNet learns representations between node pairs by considering all relations along paths, enhancing prediction accuracy and interpretability. This allows visualization of influential paths and facilitates biological validation. BioPathNet leverages a background regulatory graph (BRG) for enhanced message passing and uses stringent negative sampling to improve precision. In evaluations across various LP tasks, such as gene function annotation, drug-disease indication, synthetic lethality, and lncRNA-mRNA interaction prediction, BioPathNet consistently outperformed shallow node embedding methods, relational graph neural networks and task-specific state-of-the-art methods, demonstrating robust performance and versatility. Our study predicts novel drug indications for diseases like acute lymphoblastic leukemia (ALL) and Alzheimers, validated by medical experts and clinical trials. We also identified new synthetic lethality gene pairs and regulatory interactions involving lncRNAs and target genes, confirmed through literature reviews. BioPathNets interpretability will enable researchers to trace prediction paths and gain molecular insights, making it a valuable tool for drug discovery, personalized medicine and biology in general.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 96%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 95%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 95%
Similar papers in this journal
- Interpretable deep learning for chromatin-informed inference of transcriptional programs driven by somatic alterations across cancers 95%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 95%
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.