Expanding the utility of sequence comparisons using data from whole genomes
Gosselin, S. P.; Fullmer, M. S.; Feng, Y. P.; Gogarten, J. P.
Show abstract
Whole genome comparisons based on Average Nucleotide Identities (ANI), and the Genome-to-genome distance calculator have risen to prominence in rapidly classifying taxa using whole genome sequences. Some implementations have even been proposed as a new standard in species classification and have become a common technique for papers describing newly sequenced genomes. However, attempts to apply whole genome divergence data to delineation of higher taxonomic units, and to phylogenetic inference have had difficulty matching those produced by more complex phylogenetics methods. We present a novel method for generating reliable and statistically supported phylogenies using established ANI techniques. For the test cases to which we applied the developed approach we obtained accurate results up to at least the family level. The developed method uses non-parametric bootstrapping to gauge reliability of inferred groups. This method offers the opportunity make use of whole-genome comparison data that is already being generated to quickly produce accurate phylogenies. Additionally, the developed ANI methodology can assist classification of higher order taxonomic groups. Significance StatementThe average nucleotide identity (ANI) measure and its iterations have come to dominate in-silico species delimitation in the past decade. Yet the problem of gene content has not been fully resolved, and attempts made to do so contain two metrics which makes interpretation difficult at times. We provide a new single based ANI metric created from the combination of genomic content and genomic identity measures. Our results show that this method can handle comparisons of genomes with divergent content or identity. Additionally, the metric can be used to create distance based phylogenetic trees that are comparable to other tree building methods, while also providing a tentative metric for categorizing organisms into higher level taxonomic classifications.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Coinfinder: Detecting Significant Associations and Dissociations in Pangenomes 96%
- Taxonomic distribution of SbmA/BacA and BacA-like antimicrobial peptide transporters suggests independent recruitment and convergent evolution in host-microbe interactions 94%
- K-mer based prediction of Clostridioides difficile relatedness and ribotypes 94%
Similar papers in this journal
Similar papers in this journal
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 95%
- Genomic characterization of a diazotrophic microbiota associated with maize aerial root mucilage 94%
- CoSMIC - A hybrid approach for large-scale, high-resolution microbial profiling of novel niches 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.