Knowledge Inclusive Machine Learning for Disease Gene Prioritisation
Gamage, C. J.; Xia, Y.; Rupasinghe, R.; Senevirathne, S.; Senanayake, D.; Malepathirana, T.; Hevapathige, A.; Corbett, M.; O'Brien, T. J.; Petrou, S.; Berkovic, S. F.; Scheffer, I. E.; Gecz, J.; Bahlo, M.; Bennett, M. F.; Halgamuge, S. K.
Show abstract
The predictive performance of machine learning models depends on the context available to them. In disease gene prioritisation, this context comprises two forms: specific context from sample-level experimental data, such as gene expression and protein-protein interaction networks, and general context from accumulated and curated biological knowledge capturing established relationships among genes, diseases, and pathways. Neither is sufficient alone: experimental data are sensitive to dataset-specific noise and lack broader biological grounding, while curated knowledge lacks the resolution required for gene-level discrimination. Consequently, most machine learning approaches relying solely on experimental data risk learning spurious correlations rather than underlying biology. Here we introduce Knowledge Inclusive Machine Learning (KIML), a paradigm that integrates both context types within a unified analytical pipeline. KIML combines experimental data with two types of general context: literature-derived representations from PubMed and structured biomedical knowledge graphs. We evaluate the approach on Developmental and Epileptic Encephalopathy and benchmark it against recent methods using publicly available datasets. Performance is assessed using temporal-split evaluation and biological evaluations, including ontology enrichment analysis. KIML consistently outperforms existing approaches, providing improved predictive accuracy and biologically meaningful insights. Furthermore, the framework generates interpretable explanations of gene prioritisation and demonstrates strong generalisability across six additional diseases.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 95%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 94%
Similar papers in this journal
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 96%
- DriveWays: A Method for Identifying Possibly Overlapping Driver Pathways in Cancer 95%
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 95%
Similar papers in this journal
- HELP: A computational framework for labelling and predicting human common and context-specific essential genes 96%
- GeneCOCOA: Detecting context-specific functions of individual genes using co-expression data 95%
- Biological networks and GWAS: comparing and combining network methods to understand the genetics of familial breast cancer susceptibility in the GENESIS study 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.