MiGenPro: A linked data workflow for phenotype-genotype prediction of microbial traits using machine learning.
Loomans, M.; Suarez-Diez, M.; Schaap, P. J.; Saccenti, E.; Koehorst, J. J.
Show abstract
The availability of microbial genomic data and the development of machine learning methods have created a unique opportunity to establish associations between genetic information and phenotypes. Here, we introduce a computational workflow for Microbial Genome Prospecting (MiGenPro) that combines phenotypic and genomic information. MiGenPro serves as a workflow for the training of machine learning models that predict microbial traits from genomes that have been annotated. Microbial genomes have been consistently annotated and features were stored in a semantic framework that is easy to query using SPARQL. The data was used to train machine learning models and successfully predicted microbial traits such as motility, Gram stain, optimal temperature range, and sporulation capabilities. To ensure robustness, a hyper parameter halving grid search was used to determine optimal parameter settings followed by a five-fold cross-validation which demonstrated consistent model performance across iterations and without overfitting. Effectiveness was further validated through comparison with existing models, showing comparable accuracy, with modest variations attributed to differences in datasets rather than methodology. Classification can be further explored using feature importance characterisation to identify biologically relevant genomic features. MiGenPro provides an easy to use interoperable workflow to build and validate models to predict phenotypes from microbes based on their annotated genome. Graphical AbstractO_ST_ABSKey MessagesC_ST_ABSO_LIMiGenPro merges phenotypic and genomic data for microbial trait prediction using machine learning. C_LIO_LIMiGenPro mitigates phenotype prediction model constraints by leveraging linked data technologies. C_LIO_LIMiGenPros FAIR design enables adaptation for various phenotypes with available training data. C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CoCoPyE: feature engineering for learning and prediction of genome quality indices 95%
- DivBrowse - interactive visualization and exploratory data analysis of variant call matrices 93%
- PathoGFAIR: a collection of FAIR and adaptable (meta)genomics workflows for (foodborne) pathogens detection and tracking 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.