Predicting Protein-encoding Gene Content in Escherichia coli Genomes
Nguyen, M.; Elmore, Z.; Ihle, C.; Moen, F. S.; Slater, A. D.; Turner, B. N.; Parrello, B.; Best, A. A.; Davis, J. J.
Show abstract
In this study, we built machine learning classifiers for predicting the presence or absence of the variable genes occurring in 10-90% of all publicly available high-quality Escherichia coli genomes. The BV-BRC genus-specific protein families were used to define orthologs across the set of genomes, and a single binary classifier was built for predicting the presence or absence of each family in each genome. Each model was built using the nucleotide k-mers from a set of 100 conserved genes as features. The resulting set of 3,259 XGBoost classifiers had a per-genome average macro F1 score of 0.944 [0.943-0.945, 95% CI]. We show that the F1 scores are stable across MLSTs, and that the trend can be recapitulated through sampling with a smaller number of core genes or diverse input genomes. Surprisingly, the presence or absence of poorly annotated proteins, including "hypothetical proteins", were easily predicted (F1 = 0.902 [0.898-0.906, 95% CI]). Models for proteins with horizontal gene transfer-related functions, including transposition- (F1 = 0.895 [0.882-0.907, 95% CI]), phage- (F1 = 0.872 [0.868-0.876, 95% CI]), and plasmid-related (F1 = 0.824 [0.814-0.834, 95% CI]) functions had slightly lower F1 scores, but were still accurate. Finally, we applied the models to a holdout set of 419 diverse E. coli genomes that were isolated from freshwater environmental sources and observed an average per-genome F1 score of 0.880 [0.876-0.883, 95% CI], demonstrating the extensibility of the models. Overall, this study provides a framework for predicting variable gene content using a limited amount of input sequence data. ImportanceHaving the ability to predict the protein-encoding gene content of a genome is important for a variety of bioinformatic tasks, including assessing genome quality, binning genomes from shotgun metagenomic assemblies, and assessing risk due to the presence of antimicrobial resistance (AMR) and other virulence genes. In this study, we built a series of binary classifiers for predicting the presence or absence of variable genes occurring in 10-90% of all publicly available E. coli genomes. Overall, the results show that a large portion of the E. coli variable gene content can be predicted with high accuracy, including genes with functions relating to horizontal gene transfer.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Coinfinder: Detecting Significant Associations and Dissociations in Pangenomes 96%
- Taxonomic distribution of SbmA/BacA and BacA-like antimicrobial peptide transporters suggests independent recruitment and convergent evolution in host-microbe interactions 96%
- A validated pangenome-scale metabolic model for the Klebsiella pneumoniae species complex 95%
Similar papers in this journal
- PlasmidHostFinder: Prediction of plasmid hosts using random forest 95%
- Species-scale genomic analysis of S. aureus genes influencing phage host range and their relationships to virulence and antibiotic resistance genes 95%
- A method to correct for local alterations in DNA copy number that bias functional genomics assays applied to antibiotic-treated bacteria 95%
Similar papers in this journal
- Genomic characterization of a diazotrophic microbiota associated with maize aerial root mucilage 94%
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 94%
- Metagenome mining and functional analysis reveal oxidized guanine DNA repair at the Lost City Hydrothermal Field 94%
Similar papers in this journal
- Database size positively correlates with the loss of species-level taxonomic resolution for the 16S rRNA and other prokaryotic marker genes 96%
- A systematic pipeline for classifying bacterial operons reveals the evolutionary landscape of biofilm machineries 95%
- Predicting Antimicrobial Resistance Using Conserved Genes 95%
Similar papers in this journal
- Linkage-based ortholog refinement in bacterial pangenomes with CLARC 96%
- Automating microbial taxonomy workflows with PHANTASM: PHylogenomic ANalyses for the TAxonomy and Systematics of Microbes 96%
- Unveiling the Microbial Realm with VEBA 2.0: A modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic, and viral multi-omics from either short- or long-read sequencing 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.