Predicting lifespan-extending chemical compounds with machine learning and biologically interpretable features
Ribeiro, C.; Farmer, C. K.; de Magalhaes, J. P.; Freitas, A. A.
Show abstract
Recently, there has been a growing interest in the development of pharmacological interventions targeting ageing, as well as on the use of machine learning for analysing ageing-related data. In this work we use machine learning methods to analyse data from DrugAge, a database of chemical compounds (including drugs) modulating lifespan in model organisms. To this end, we created four datasets for predicting whether or not a compound extends the lifespan of C. elegans (the most frequent model organism in DrugAge), using four different types of predictive biological features, based on compound-protein interactions, interactions between compounds and proteins encoded by ageing-related genes, and two types of terms annotated for proteins targeted by the compounds, namely Gene Ontology (GO) terms and physiology terms from the WormBases Phenotype Ontology. To analyse these datasets we used a combination of feature selection methods in a data pre-processing phase and the well-established random forest algorithm for learning predictive models from the selected features. The two best models were learned using GO terms and protein interactors as features, with predictive accuracies of about 82% and 80%, respectively. In addition, we interpreted the most important features in those two best models in light of the biology of ageing, and we also predicted the most promising novel compounds for extending lifespan from a list of previously unlabelled compounds.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GenEpi: Gene-based Epistasis Discovery Using Machine Learning 94%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 94%
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 94%
Similar papers in this journal
- A cautionary tale about properly vetting datasets used in supervised learning predicting metabolic pathway involvement 95%
- Two-step multi-omics modelling of drug sensitivity in cancer cell lines to identify driving mechanisms 94%
- Automated recognition of functional compound-protein relationships in literature 94%
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 94%
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 94%
- VICTOR: A visual analytics web application for comparing cluster sets 94%
Similar papers in this journal
Similar papers in this journal
- Deep learning approach for automatic assessment of schizophrenia and bipolar disorder in patients using R-R intervals 94%
- Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathways 94%
- Novel feature selection methods for construction of accurate epigenetic clocks 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.