Machine Learning-Driven Phenotype Predictions based on Genome Annotations
Edirisinghe, J. N.; Goyal, S.; Brace, A.; Colasanti, R.; Gu, T.; Sadhkin, B.; Zhang, Q.; Kamimura, R.; Henry, C. S.
Show abstract
Over the past two decades, there has been a remarkable and exponential expansion in the availability of genome sequences, encompassing a vast number of isolate genomes, amounting to hundreds of thousands, and now extending to millions of metagenome-assembled genomes. The rapid and accurate interpretation of this data, along with the profiling of diverse phenotypes such as respiration type, antimicrobial resistance, or carbon utilization, is essential for a wide range of medical and research applications. Here, we leverage sequenced-based functional annotations obtained from the RAST annotation algorithm as predictors and employ six machine learning algorithms (K-Nearest Neighbors, Gaussian Naive Bayes, Support Vector Machines, Neural Networks, Logistic Regression, and Decision Trees) to generate classifiers that can accurately predict phenotypes of unclassified bacterial organisms. We apply this approach in two case studies focused on respiration types (aerobic, anaerobic, and facultative anaerobic) and Gram-stain types (Gram negative and Gram positive). We demonstrate that all six classifiers accurately classify the phenotypes of Gram stain and respiration type, and discuss the biological significance of the predicted outcomes. We also present four new applications that have been deployed in The Department of Energy Systems Biology Knowledgebase (KBase) that enable users to: (i) Upload high-quality data to train classifiers; (ii) Annotate genomes in the training set with the RAST annotation algorithm; (iii) Build six different genome classifiers; and (iv) Predict the phenotype of unclassified genomes. (https://narrative.kbase.us/#catalog/modules/kb_genomeclassification)
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- No one tool to rule them all: Prokaryotic gene prediction tool performance is highly dependent on the organism of study 95%
- Reconstructor: A COBRApy compatible tool for automated genome-scale metabolic network reconstruction with parsimonious flux-based gap-filling 94%
- DeepES: Deep learning-based enzyme screening to identify orphan enzyme genes 93%
Similar papers in this journal
- FeGenie: a comprehensive tool for the identification of iron genes and iron gene neighborhoods in genomes and metagenome assemblies 94%
- No Assembly Required: Using BTyper3 to Assess the Congruency of a Proposed Taxonomic Framework for the Bacillus cereus group with Historical Typing Methods 94%
- ProkBERT Family: Genomic Language Models for Microbiome Applications 93%
Similar papers in this journal
- NFixDB (Nitrogen Fixation DataBase) - A Comprehensive Integrated Database for Robust 'Omics Analysis of Diazotrophs 94%
- BacTermFinder: A Comprehensive and General Bacterial Terminator Finder using a CNN Ensemble 94%
- Life at the extremes: Maximally divergent microbes with similar genomic signatures linked to extreme environments 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.