The effects of biological knowledge graph topology on embedding-based link prediction
Bradshaw, M. S.; Avramov, A.; Gaskell, A.; Layer, R. M.
Show abstract
Due to the limited information available about rare diseases and their causal variants, knowledge graphs are often used to augment our understanding and make inferences about new gene-disease connections. Knowledge graph embedding methods have been successfully applied to various biomedical link prediction tasks but have yet to be adopted for rare disease variant prioritization. Here, we explore the effect of knowledge graph topology on knowledge graph embedding link prediction performance and challenge the assumption that massively aggregating knowledge graphs is beneficial in deciphering rare disease cases and improving prediction outcomes. We find that using a filtered version of the Monarch knowledge graph with only 11% of the original size results in notably improved model predictive performance. Additionally, these findings suggest that successful KG optimization depends on selecting high-quality information rather than simply maximizing the amount of data included.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting candidate genes from phenotypes, functions, and anatomical site of expression 96%
- DTI-Voodoo: machine learning over interaction networks and ontology-based background knowledge predicts drug-target interactions 95%
- RAPPPID: Towards Generalisable Protein Interaction Prediction with AWD-LSTM Twin Networks 95%
Similar papers in this journal
Similar papers in this journal
- HARVESTMAN: A framework for hierarchical featurelearning and selection from whole genome sequencingdata 95%
- SKiM-GPT: Combining Biomedical Literature-Based Discovery with Large Language Model Hypothesis Evaluation 94%
- Optimal construction of a functional interaction network from pooled library CRISPR fitness screens 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.