A new method for rapid genome classification, clustering, visualization, and novel taxa discovery from metagenome
Wang, Z.; Ho, H.; Egan, R.; Yao, S.; Kang, D.; Froula, J.; Sevim, V.; Schulz, F.; Shay, J. E.; Macklin, D.; McCue, K.; Orsini, R.; Barich, D. J.; Sedlacek, C. J.; Li, W.; Morgan-Kiss, R.; Woyke, T.; Slonczewski, J. L.
Show abstract
Current supervised phylogeny-based methods fall short on recognizing species assembled from metagenomic datasets from under-investigated habitats, as they are often incomplete or lack closely known relatives. Here, we report an efficient software suite, \"Genome Constellation\", that estimates similarities between genomes based on their k-mer matches, and subsequently uses these similarities for classification, clustering, and visualization. The clusters of reference genomes formed by Genome Constellation closely resemble known phylogenetic relationships while simultaneously revealing unexpected connections. In a dataset containing 1,693 draft genomes assembled from the Antarctic lake communities where only 40% could be placed in a phylogenetic tree, Genome Constellation improves taxa assignment to 61%. It revealed six clusters derived from new bacterial phyla and 63 new giant viruses, 3 of which missed by the traditional marker-based approach. In summary, we demonstrate that Genome Constellation can tackle the computational and algorithmic challenges in large-scale taxonomy analyses in metagenomics.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Ultra-accurate Microbial Amplicon Sequencing with Synthetic Long Reads 94%
- METABOLIC: High-throughput profiling of microbial genomes for functional traits, biogeochemistry, and community-scale metabolic networks 94%
- MetaEuk - sensitive, high-throughput gene discovery and annotation for large-scale eukaryotic metagenomics 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.