Biological machine learning combined with bacterial population genomics reveals common and rare allelic variants of genes to cause disease
Bandoy, D. D. R.; Weimer, B. C.
Show abstract
Highly dimensional data generated from bacterial whole genome sequencing is providing unprecedented scale of information that requires appropriate statistical frameworks of analysis to infer biological function from bacterial genomic populations. Application of genome wide association study (GWAS) methods is an emerging approach with bacterial population genomics that yields a list of genes associated with a phenotype with an undefined importance among the candidates in the list. Here, we validate the combination of GWAS, machine learning, and pathogenic bacterial population genomics as a novel scheme to identify SNPs and rank allelic variants to determine associations for accurate estimation of disease phenotype. This approach parsed a dataset of 1.2 million SNPs that resulted in a ranked importance of associated alleles of Campylobacter jejuni porA using multiple spatial locations over a 30-year period. We validated this approach using previously proven laboratory experimental alleles from an in vivo guinea pig abortion model. This approach, termed BioML, defined intestinal and extraintestinal groups that have differential allelic variants that cause abortion. Divergent variants containing indels that defeated gene callers were rescued using biological context and knowledge that resulted in defining rare and divergent variants that were maintained in the population over two continents and 30 years. This study defines the capability of machine learning coupled to GWAS and population genomics to simultaneously identify and rank alleles to define their role in abortion, and more broadly infectious disease.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- K-mer based prediction of Clostridioides difficile relatedness and ribotypes 93%
- Comparison of gene-by-gene and genome-wide short nucleotide sequence based approaches to define the global population structure of Streptococcus pneumoniae 93%
- SNPPar: identifying convergent evolution and other homoplasies from microbial whole-genome alignments 93%
Similar papers in this journal
- Identifying the essential genes of Mycobacterium avium subsp. hominissuis with Tn-Seq using a rank-based filter procedure. 92%
- Genomic epidemiology of vancomycin resistant Enterococcus faecium (VREfm) in Latin America: Revisiting the global VRE population structure 92%
- Deciphering the metabolic capabilities of Bifidobacteria using genome-scale metabolic models 92%
Similar papers in this journal
- Cov2clusters: genomic clustering of SARS-CoV-2 sequences 92%
- DNA Extraction Method Optimized for Nontuberculous Mycobacteria Long-Read Whole Genome Sequencing 91%
- GeneMates: an R package for Detecting Horizontal Gene Co-transfer between Bacteria Using Gene-gene Associations Controlled for Population Structure 91%
Similar papers in this journal
- Population structure and pangenome analysis of Enterobacter bugandensis uncover the presence of blaCTX-M-55, blaNDM-5 and blaIMI-1, along with sophisticated iron acquisition strategies 91%
- Origin and evolutionary dynamics of multi-drug resistant and highly virulent community-associated methicillin-resistant Staphylococcus aureus ST772-SCCmec V lineage 90%
- Informing plasmid compatibility with bacterial hosts using protein-protein interaction data 90%
Similar papers in this journal
- Convergence of resistance and evolutionary responses in Escherichia coli and Salmonella enterica co-inhabiting chicken farms in China 95%
- A comprehensive update to the Mycobacterium tuberculosis H37Rv reference genome 93%
- Construction of a complete set of Neisseria meningitidis defined mutants - the NeMeSys 2.0 collection - and its use for the phenotypic profiling of the genome of an important human pathogen 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.