Ontology pre-training improves machine learning-based predictions for metabolites
Tumescheit, C.; Glauer, M.; Fluegel, S.; Larralde, M.; Neuhaus, F.; Mossakowski, T.; Hastings, J.
Show abstract
Recent advances in the field of machine learning have shown that integration of expert knowledge improves performances, in particular for complex domains such as biology. Bio-ontologies offer a rich source of curated biological knowledge that can be harnessed to this end. Here, we describe an intuitive and generalisable approach to embed the knowledge contained in a classification hierarchy derived from a bio-ontology into a machine learning model as an intermediate training step between general-purpose pre-training and task-specific fine-tuning in a process that we call ontology pre-training. We show that this approach leads to an improvement in predictive performance and a reduction in training time for a broad range of prediction tasks relevant to understanding metabolite functions in living systems, using a range of datasets derived from MoleculeNet. We see the biggest improvement for regression tasks, e.g. prediction of lipophilicity and aqueous solubility of molecules, and a robust improvement for most classification tasks. Our approach can be adapted for a wide range of knowledge sources, models and prediction tasks.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluation of network architecture and data augmentation methods for deep learning in chemogenomics 97%
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 97%
- DeepGraphMol, a multi-objective, computational strategy for generating molecules with desirable properties: a graph convolution and reinforcement learning approach 96%
Similar papers in this journal
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 96%
- A Graph-Attention-Based Deep Learning Network for Predicting Biotech-Small-Molecule Drug Interactions 96%
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.